Anthropic settles the copyright lawsuit with the highest amount ever of $1.5 billion: but the authors and the publisher don't think they won

A U.S. federal judge approved the final settlement of Anthropic’s copyright class action lawsuit, with the amount reaching $1.5 billion—setting the highest settlement record in U.S. copyright law history.
(Background: After Anthropic scanned 2 million books to train Claude, it directly fed them into a shredder to destroy the books.)
(Additional context: Sony goes to trial this month against Suno—whether using copyright songs to train AI counts as “fair use” will be argued in court for the first time.)

Table of contents

Toggle

  • A trial about “how to get the books”
  • The authors don’t think they won
  • The battlefield hasn’t ended—just changed opponents

After more than a year of litigation, the case has finally reached the step where Anthropic pays to put an end to it. On the 20th, a U.S. federal judge formally approved the final settlement of this copyright class action. About 500k works were found to be involved, with an average compensation of $3,000 per work. The payment will be split between authors and publishers who hold the rights, totaling $1.5 billion.

This final approval was signed by U.S. District Judge Araceli Martinez-Olguin of the Northern District of California, taking over from the preliminary ruling issued last year by Judge William Alsup. The case appears to be over on the surface, but the controversy left by Alsup’s decision back then is actually only just starting to spread across the entire AI industry.

A trial about “how to get the books”

The story goes back to 2024, when a group of authors sued Anthropic, alleging it used pirated books to train the chat bot Claude without authorization.

During the proceedings, Alsup issued a rare two-sided ruling: he found that training an AI model on copyrighted text constitutes fair use. In simple terms, feeding books to a model to learn language rules is not, by itself, copyright infringement; legally, it is treated as a transformative use similar to citation, commentary, or educational purposes.

At the time, this finding was widely seen as a turning point for the AI industry—effectively opening a door to the legality of training data, and giving a glimpse of hope to other AI companies still fighting in court.

But Alsup also drew another line: just because the training itself is lawful doesn’t mean the method of obtaining the books is also lawful. Anthropic obtained books through two channels. One was buying them with money and then scanning them—there was no dispute. The other was downloading directly from pirated websites such as Library Genesis and Pirate Library Mirror. Alsup ruled that this portion was illegal and said the issue could go straight to a jury trial.

Once it went to trial, the amount of damages would be left to the jury’s discretion, and the ceiling could far exceed the settlement figure. After weighing its options, Anthropic chose a settlement amount it could predict, trading off a result that was hard to forecast and a dispute that would also take longer.

This decision looks pragmatic, yet it also ensures the public will never see in the official judgment what price the jury would have put on each pirated book.

The authors don’t think they won

$1.5 billion sounds astronomical, but many authors do not view this settlement as a victory. The reason is that the money compensates the “pirated acquisition” side path—not the “whether AI training is infringement” main path.

The central legal question—whether training models with copyrighted works counts as fair use—actually produced a result favorable to AI companies. What the authors got was essentially a penalty for Anthropic being “lazy” and taking a shortcut to avoid paying for books, not a principled win for “AI companies can’t use my books to train models.” In other words, the authors won the least important battle, but lost the most critical one in the war.

More importantly, this ruling doesn’t even rise to the level of binding precedent. Plainly put, later judges handling similar cases do not need to copy Alsup’s logic; they can make entirely different conclusions based on the facts of the case. The reason is simple: this is only a decision from a “single district court.” And because Anthropic chose to settle, the case will never be appealed to a higher court, so there is no chance for a higher-level court to confirm and expand the rule into one the whole industry must follow.

A district-court ruling versus a binding precedent for the entire industry—there’s a whole appeals process in between. And this settlement conveniently blocks that entire path, sending the question of “whether AI training is fair use” back into the legal gray zone.

For other AI companies still suing, Alsup’s ruling can at most be used as a referenced case—no judge is obligated to write along the same lines.

The battlefield hasn’t ended—just changed opponents

The end of this Anthropic case doesn’t mean the copyright dispute over AI training data is over. Google, Meta, Midjourney, and OpenAI all still have a string of copyright lawsuits. Each case has different circumstances; the evidence in the hands of the judges and the ways the data was obtained also differ. In the future, courts could very well reach conclusions completely different from Alsup’s.

Just last week, publishers including Hachette, Cengage, and Elsevier, along with author Scott Turow and S.C.R.I.B.E., filed a class action lawsuit against Google, accusing it of using copyrighted works to train the Gemini AI platform. Same script, different defendant—the story seems likely to repeat, and the next judge doesn’t have to refer to a single word from the Anthropic decision.

For the publishing industry as a whole, the Anthropic case is more like a template, demonstrating that the path of “conceding on fair use first, then negotiating a settlement amount based on illegal pirated acquisition” can work. That also implies that in future lawsuits against Google or Meta, the battlefield will likely focus on the method of obtaining data—not on whether training itself is legal. The difference is that each company’s training-data source records are not the same: some claim licensed partnerships, while others can’t explain the source clearly. Those details will determine whether the next lawsuit ends with the same settlement outcome—or whether it gets sent to court and fought out all the way.

View Original
This page may contain third-party content, which is provided for information purposes only (not representations/warranties) and should not be considered as an endorsement of its views by Gate, nor as financial or professional advice. See Disclaimer for details.
  • Reward
  • Comment
  • Repost
  • Share
Comment
Add a comment
Add a comment
No comments
  • Pinned