Tech

Anthropic's landmark $1.5bn AI copyright settlement gets final approval

TechCrunch9 h ago
A judge's gavel resting on a stack of law books
A judge's gavel resting on a stack of law booksPhoto: KATRIN BOLOVTSOVA / Pexels

A US federal court has granted final approval to Anthropic's $1.5 billion settlement resolving claims that the AI company used copyrighted books without permission to train its Claude models, bringing to a close one of the most closely watched legal disputes to emerge from the generative AI boom. The settlement, first announced last year, is among the largest publicly disclosed payouts by an AI company over training-data claims to date.

The underlying case centered on allegations, brought by a group of authors, that Anthropic had used pirated or otherwise unauthorized copies of books, including works obtained from so-called shadow libraries that host copyrighted material without licensing agreements, as part of the massive text datasets used to train its large language models. The plaintiffs argued this practice amounted to copyright infringement at a scale affecting potentially hundreds of thousands of individual works.

Under the terms of the settlement, Anthropic agreed to pay the $1.5 billion sum to compensate authors and rights holders whose works were used without authorization, while not admitting wrongdoing as part of the deal, a common structure in large corporate settlements that allows a company to resolve litigation without a formal court finding of liability. The company has said the settlement allows it to focus on building its products rather than prolonged litigation.

A federal judge overseeing the case had earlier issued a preliminary ruling that drew significant attention across the AI industry, distinguishing between the use of legally acquired books for AI training, which the judge suggested could potentially qualify as fair use, and the use of pirated copies, which the ruling characterized far more skeptically. That distinction has been widely cited by legal analysts as an early signal of how courts might approach the broader wave of AI copyright litigation still working through the system.

The case is one of several major lawsuits filed against AI companies by authors, publishers, news organizations and other content creators since the launch of ChatGPT and similar tools triggered a wave of legal challenges over how large language models are trained. Other prominent disputes, including cases involving OpenAI, Meta and Stability AI, remain unresolved and are working through courts in parallel, each testing slightly different legal theories about fair use, market harm and the specific ways training data was obtained.

Industry analysts note that a settlement of this size, while enormous in absolute terms, is unlikely to be viewed by AI companies as a deterrent significant enough to change training practices broadly, given the scale of revenue and investment flowing into the sector. Some legal observers argue the more consequential outcomes will come from ongoing cases that produce binding appellate rulings on the underlying fair-use question, rather than settlements that resolve individual disputes without setting broad legal precedent.

Author advocacy groups involved in the litigation described the settlement as an important acknowledgment that AI companies cannot treat copyrighted creative work as freely available training material, while also cautioning that a cash payment does not by itself establish a durable legal framework governing how AI companies should license or compensate for the use of copyrighted material going forward.

Anthropic, alongside other major AI labs, has in parallel been pursuing licensing agreements directly with publishers and content owners, a trend industry watchers see as an attempt to preempt further litigation by establishing negotiated terms for data use before disputes reach court. The scale and terms of such deals vary considerably and are rarely disclosed in full detail.

The settlement does not resolve broader regulatory or legislative questions about AI and copyright that remain active in multiple jurisdictions, including ongoing debates in the US Congress and various international bodies over whether existing copyright law adequately addresses the use of creative works in training machine learning systems, or whether new legislation is needed.

For now, the approval closes a significant chapter in Anthropic's legal history, but the broader legal landscape around AI training data remains unsettled, with courts, legislators and the AI industry itself still working out what rules, if any, should govern how models are built from the vast troves of text, images and other creative material available online.

This article is an AI-curated summary based on TechCrunch. The illustration is a stock photo by KATRIN BOLOVTSOVA from Pexels.

Read next