Publishers sue Google over Gemini AI training, calling it one of history's biggest copyright violations
Hachette, Cengage, and Elsevier claim Google used millions of copyrighted books without permission to build its Gemini AI. A federal lawsuit filed in New York could reshape how AI companies gather training data.

Key points
- Three major publishers and bestselling author Scott Turow filed a federal lawsuit against Google in New York.
- The plaintiffs are Hachette Book Group, Cengage Learning, and Elsevier, covering fiction, education, and academic publishing.
- The suit accuses Google of using millions of copyrighted books without permission to train Gemini, its family of AI models.
- The plaintiffs call it "one of the most prolific infringements of copyrighted materials in history."
- The case adds to a growing wave of copyright litigation targeting large AI companies over how they build their systems.
Three major book publishers and a celebrated American novelist are taking Google to federal court, accusing the tech giant of stealing millions of books to teach its artificial intelligence how to write, reason, and answer questions.
Hachette Book Group, Cengage Learning, and Elsevier filed the suit in New York alongside author Scott Turow, who has sold millions of legal thrillers. The case, first reported by The Guardian, targets Google's Gemini, the company's family of large language models, the technology behind AI assistants like Google's own AI Overview search feature and chatbot.
The publishers allege Google copied their books on a massive scale and fed them into Gemini's training process, the stage where an AI system reads enormous amounts of text to learn patterns in language. Training on books is especially valuable because books contain long, carefully constructed arguments and prose, qualities that help an AI respond coherently.
None of that copying was authorised, the publishers say.
"One of the most prolific infringements of copyrighted materials in history" is the phrase used in the complaint. That is a striking claim, but it fits a pattern. OpenAI, Meta, and other AI developers face similar lawsuits from authors, news organisations, and other rights holders who argue their work was taken without consent or payment.
What does this mean for people who read or buy books?
For now, nothing changes on your shelf or your e-reader. This is a civil dispute over how Google built its technology, not a product recall. But the outcome could matter quite a lot down the line.
If courts decide AI companies must licence the books they train on, publishers and authors gain a new revenue stream. Prices for AI products could rise to cover those costs, or companies might train on smaller, licensed datasets, which could affect how capable the next generation of AI tools turns out to be.
The three publishers between them cover a wide range of reading and learning. Hachette publishes popular fiction and non-fiction. Cengage focuses on educational textbooks used in schools and colleges. Elsevier publishes scientific and medical research journals. The breadth of that coalition signals the lawsuit is not a niche complaint from one corner of publishing.
Google has not yet filed a public response to the complaint. The company has previously argued, in other copyright cases, that training AI on publicly available text falls under "fair use," a legal doctrine that allows limited use of copyrighted material without permission under certain conditions.
Federal courts have not yet settled the question definitively. That makes this case, and others like it, ones worth watching closely.



