NEW YORK / RankWire.AI / – Hachette Book Group, Cengage Learning and Elsevier have initiated legal action against Google concerning its Gemini artificial intelligence platform. Author Scott Turow and his organization, S.C.R.I.B.E., have joined the class action lawsuit. The suit was filed on July 10 in the U.S. District Court for the Southern District of New York. The plaintiffs accuse Google of copying millions of copyrighted books and journal articles without authorization during the development and training of Gemini. As of July 15, the court had yet to rule on the allegations or certify the proposed class.

According to the complaint, Google sourced material via Google Books, Google Play Books, and Google Scholar. Publishers and authors had provided works for specific functionalities such as search, sales, and research access. The plaintiffs argue that these arrangements did not permit extensive commercial AI training. They also claim Google downloaded sizable web-scraped datasets containing copyrighted content. The document states that some of this material originated from known pirate sources and services protected by paywalls.
The 57-page complaint outlines four claims under federal law. Three relate to alleged reproduction via Google services, web scraping, and Gemini’s development or training. The fourth claim invokes the Digital Millennium Copyright Act, alleging Google removed or altered copyright management information from training data. The filing also references internal discussions about using publisher-supplied books. One assessment cited estimates potential fines between $10 billion and $100 billion. These allegations have yet to be tested in court.
Class Definition Encompasses Registered Works
The proposed class includes owners of registered U.S. copyrights for qualifying books and journal articles. To qualify, books must have an International Standard Book Number (ISBN), while articles need a Digital Object Identifier (DOI) or International Standard Serial Number (ISSN). The class covers works allegedly copied from Google services or obtained through web scraping, as well as those reproduced during Gemini’s development or training phases.
Eligibility also depends on registration timing. One criterion requires registration within five years of publication and prior to Google’s alleged reproduction or distribution. Another requires registration within three months of publication. The suit excludes government entities, Google affiliates, certain court participants, and individuals who properly opt out. The court must approve the class designation before the case advances to a broader group.
Legal Claims for Damages and Transparency
The plaintiffs seek either statutory damages or actual damages for proven infringements. They also request Google’s profits attributable to any confirmed copyright violations. Their remedies include an injunction, legal costs, and a jury trial. The complaint does not specify a total damages figure. It demands that Google disclose Gemini training data, data collection methods, and known model capabilities through a court-ordered accounting.
This accounting would identify copyrighted works used to train Gemini and detail how Google collected, copied, processed, and encoded these materials. The plaintiffs also seek court-supervised destruction of unauthorized copies in Google’s possession. Earlier, Hachette and Cengage aimed to join separate Google generative AI litigation in California. The New York case now includes Elsevier, Turow, and S.C.R.I.B.E., focusing on claims related to Google services, web scraping, and Gemini’s training process.
