A newly filed copyright lawsuit against Google quotes what it says is an internal company document warning that training Gemini on Play Books material was “highly problematic for Google,” with financial exposure the document allegedly pegs at “$10Bs-$100Bs.” Three publishers and novelist Scott Turow are the plaintiffs behind that claim, and the case turns on whether content supplied to Google for one purpose can lawfully become training data for another. The document’s language is an allegation drawn from the plaintiffs’ complaint, not a disclosed Google estimate, and it has not been tested in court.

Hachette Book Group, Cengage Learning, Elsevier, Turow, and his company S.C.R.I.B.E. filed the proposed class action on July 10, 2026, in the U.S. District Court for the Southern District of New York. The Association of American Publishers announced the filing the same day. The suit contends that Google took material submitted for search indexing and retail sales through Google Books, Play Books, and Google Scholar, and repurposed it to train a commercial AI model without separate authorization.

The complaint raises four claims. Three target different sourcing paths under federal copyright protections: material drawn from Google Books, Play Books, and Scholar; separate copies the suit says Google pulled from pirated and paywalled sources during web scraping; and copying plaintiffs allege happened again inside the training process itself. A fourth claim, brought under the DMCA, accuses Google of stripping copyright management information the works originally carried. None of the four claims have been adjudicated, and Google had not issued a public response as of this report.

Beyond damages and an injunction, the plaintiffs want a full accounting of every source Gemini’s training data came from, plus court orders forcing deletion of any copies obtained without authorization. That remedy matters more than the headline dollar figure. AI developers have so far kept training-data provenance almost entirely opaque, disclosing categories of sources in policy papers but not verifiable inventories.

A discovery order compelling that kind of inventory would give publishers and SEO teams something they currently lack: a verified record of which specific titles and domains fed a frontier model, rather than an inference built from crawl logs and robots.txt compliance. That distinction would let licensing teams price access to specific content categories instead of negotiating blind, and it would give SEOs advising publisher clients a factual basis for AI-training risk assessments they now have to guess at.

Neither sourcing path in the complaint runs through Google-Extended, the robots.txt token that governs whether Google can use crawled content for Gemini training and grounding. The books arrived under direct supply agreements, and the alleged scraped copies trace to pirate and subscription sites publishers never controlled, so no crawler directive would have stopped either route. Google’s own June 25 policy paper defends training on public web data as fair use while pointing to Google-Extended as an opt-out mechanism, a position this case does not directly test.

Publishers should not treat robots.txt settings as a complete defense against AI-training exposure, since this case alleges Google obtained material through supply agreements and third-party scrapes that crawler directives never reach. Anyone who has supplied content to a Google product under a narrow-use agreement should document those original terms now, because that paper trail, not the site’s current crawl settings, is what a court will examine first.

Search Engine Journal’s Matt G. Southern first reported the lawsuit and its internal-document allegations on July 17, 2026.