copyright

When a Novelist's Life's Work Becomes a Training Token

Scott Turow and major publishers sued Google, alleging it fed millions of copyrighted books into Gemini despite an internal note warning of billions in fines.

Scott Turow has spent forty years writing books that people carry onto airplanes and finish in a single sitting. Now his novels sit inside a lawsuit alleging that Google fed them, and millions of other copyrighted works, into the machine that powers Gemini. The claim at the heart of the case is simple and unsettling: the labor of making culture was quietly converted into raw material for a product the makers never agreed to build.

Turow is not litigating alone. Hachette, Cengage, and Elsevier joined him in a class action filed in the Southern District of New York, arguing that Google copied books and journal articles wholesale and, according to the original report, stripped out copyright management information to obscure what it had taken. The most damning detail in the filing is not a statistic but a sentence Google allegedly wrote to itself: an internal note warning that training on such books was 'highly problematic,' carrying '$10Bs-$100Bs in potential fines.'

The training happened anyway. That is the part worth sitting with. A company did the math on exposure measured in the tens of billions, wrote down that the risk was severe, and proceeded — a wager that the value of a smarter model outweighed the cost of being caught. Whether or not the court agrees, the internal note reframes the whole debate: this was not an accident of scale or a gray area misread. It reads, in the complaint's telling, like a decision.

What does a writer actually lose here?

What a novelist loses in this arrangement is not only royalties. It is standing. When a book becomes a training token, the author no longer has a seat at the table where the terms are set — no ability to say yes for a fee, no ability to say no on principle, no way to know their work was even in the room. The publishers can sue because they have lawyers and balance sheets. The writer who self-published a memoir, or the mid-list author whose contract predates the word 'AI,' has neither. The bargaining position that copyright was supposed to guarantee simply evaporates at the moment of ingestion.

There is a widening rift here between the people who make culture and the systems that consume it, and the Turow suit is one of the first times a working novelist has stood on the litigation side of it rather than the abstract one. For years the conversation about AI and creativity stayed philosophical — will machines write novels, should they. The lawsuits have made it concrete. The question is no longer whether a model can imitate a voice. It is whether the voice's owner had any say in teaching it to.

The legal argument will turn on fair use, on the removal of copyright information, on the mechanics of how training data is assembled. But the human argument is about consent, and consent is the thing that scales badly. Google can ask a million questions of a model in a second; it cannot, apparently, ask a million authors for permission without the economics collapsing. That asymmetry is the actual story. The technology that makes it trivial to reproduce human expression also makes it inconvenient to honor the person who produced it.

I keep returning to that internal note. A company can know, in writing, that a thing is wrong and expensive and do it anyway, because the upside is a better product and the downside is a line item. Culture, for the people who make it, is not a line item. It is the forty years Turow spent at the desk. The lawsuit will be decided on precedent and statute. But the rift it exposes — between the ones who write and the ones who harvest — will not close in a courtroom.

FAQ

What are the publishers and Scott Turow actually alleging against Google?

The class action, filed in the Southern District of New York, alleges that Google copied millions of copyrighted books and journal articles to train Gemini and removed copyright management information to conceal the copying. The complaint cites an internal Google note describing the practice as 'highly problematic,' with potential fines in the tens to hundreds of billions of dollars.

Why does this lawsuit matter beyond the parties involved?

The suit marks one of the first times a working novelist, alongside major publishers, has formally challenged how AI models are trained on copyrighted work. It reframes the debate over AI and creativity from a philosophical question into a concrete fight over consent, compensation, and whether authors have any say when their work becomes training data.