NEW YORK / RankWire.AI / – Three leading U.S. publishing companies have filed a lawsuit against Google, claiming copyright infringement related to its Gemini artificial intelligence platform. Hachette Book Group, Cengage Learning, and Elsevier initiated a class action alongside author Scott Turow and his firm, S.C.R.I.B.E. The complaint was submitted on July 10 in a federal court in New York, alleging that Google copied millions of copyrighted books and journal articles without authorization during the training and development of Gemini models.

The plaintiffs contend that Google acquired the material via Google Books, Google Play Books, and Google Scholar. Publishers and authors had provided their works to facilitate search, sales, and research functionalities, according to the complaint. However, the filing states these arrangements did not permit Google to reproduce the works for commercial AI training purposes. It further accuses Google of utilizing web-scraped datasets that included content from pirate websites and subscription services protected behind paywalls.
Google faces four allegations detailed in the 57-page complaint. Three of these claims concern unauthorized reproduction through Google services, web scraping, and the development or training of Gemini. The fourth invokes the Digital Millennium Copyright Act, asserting that Google removed or altered copyright management information, including author names, ownership data, and publication details. As of July 15, the court had yet to rule on these claims or grant class-action status.
Four Allegations Focus on Gemini’s Training Data
The proposed class encompasses owners of registered U.S. copyrights in books and journal articles. To qualify, books must have an International Standard Book Number, while articles need a Digital Object Identifier or an International Standard Serial Number. This definition pertains to works Google allegedly copied from its platforms, downloaded via web scraping, or reproduced during Gemini’s development. Eligibility is limited to works registered within the deadlines specified in the complaint.
Examples of the alleged copying include works from Hachette, Cengage, and Elsevier. The categories affected range from fiction and textbooks to scholarly articles. The complaint also references internal Google evaluations concerning potential legal risks associated with publisher-provided works. One internal assessment reportedly warned of possible fines between $10 billion and $100 billion, though the court has not made any findings regarding these internal documents.
Plaintiffs Demand Damages and Transparency
The plaintiffs are seeking statutory damages or actual damages along with profits from any proven infringement. They also request an injunction, recovery of legal expenses, and a jury trial. Their proposed order would compel Google to disclose the materials and methods used to train Gemini. Additionally, they ask the court to oversee the destruction of any unauthorized copies under Google’s control. The complaint does not specify an exact amount for damages sought.
This New York case follows an earlier effort by Hachette and Cengage to join separate copyright litigation involving Google AI in California. The Association of American Publishers noted that this new case preserves certain claims outside the scope of that proceeding. The current lawsuit includes Elsevier, Turow, and S.C.R.I.B.E., alongside the two publishers, and requests the New York court to determine if Google’s Gemini training and data collection efforts violate federal copyright law and the Digital Millennium Copyright Act.
