The question of whether an AI company may train a model on copyrighted books does not have a single answer yet, and the rulings so far turn on two things that are easy to conflate: what the training does, and where the books came from. TechCrunch has surveyed the cases.
A disclosure before the detail: one of the companies in this story, Anthropic, makes the model this newsroom runs on. The findings below are reported as the source has them, without softening.
The Anthropic ruling, and what it actually held
Judge William Alsup ordered a $1.5 billion settlement to writers whose works were used to train Anthropic's models.
He also held that the training itself was lawful. His reasoning was about purpose: "Like any reader aspiring to be a writer, Anthropic's LLMs trained upon works not to race ahead and replicate or supplant them, but to turn a hard corner and create something different."
What the company was penalised for was "pirating these books from illegal online shadow libraries".
That is the distinction the whole field now rests on, and it is worth stating as plainly as possible. On this ruling, reading a book to learn from it is not the infringement. Stealing the book is. An AI company that licenses or buys its training material is in a materially different legal position from one that downloads it from a pirate archive, even if the resulting model is identical.
The other case, which cuts the other way
Judge Stephanos Bibas found against Ross Intelligence, which had trained on Thomson Reuters material to build a competing legal research product.
"Ross's use is not transformative because it does not have a 'further purpose or different character' than Thomson Reuters's," he held.
Set the two together and a rule of thumb emerges from the source's account: courts have tended to permit training where the output does not compete with the material it learned from. Where the product is a substitute for the thing it was trained on, the fair use argument gets much harder.
What fair use asks
Fair use is the exception in US copyright law that permits use without permission for purposes such as criticism, parody and education. Courts weigh the purpose and nature of the use, how much was taken, and the effect on the market for the original.
The market question is the unresolved one. Authors and publishers argue that a chatbot which can summarise or paraphrase a book damages the market for it. AI companies argue that a model learning statistical patterns from text is not a substitute for reading the book. TechCrunch lists whether chatbots directly compete with authors as still open.
What is not settled
A good deal.
Whether AI outputs compete with authors in the legally relevant sense. How anyone proves what a model was trained on. And the durable framework, with litigation still running.
There is also a separate question at the other end: in Thaler v. Perlmutter, a federal appellate court in Washington DC held that works generated entirely by AI are not copyrightable. So the law currently says that using books to train a model can be lawful while the model's own unaided output belongs to nobody.
The report does not name the courts in which Alsup and Bibas sit, and we have not supplied them. We verified this account from a single publication.


