Courts and lawyers continue to grapple with whether using copyrighted books to train large language models is permitted under current law. High-profile decisions have produced divergent outcomes that hinge on source material, purpose and how courts interpret a copyright statute last revised in 1976.
Alsup ruling: lawful training, penalty for pirated sources
In one early landmark decision, Judge William Alsup ruled that training an AI on published works can be lawful but still sanctioned Anthropic with a $1.5 billion settlement tied to the company’s use of books obtained from illegal online shadow libraries. Alsup compared an LLM’s ingestion of vast text corpora to a writer studying literature, and distinguished training from straightforward copying.
Copyright attorney Cathy Gellis said the ruling favors AI developers because the judge treated training as analogous to reading rather than copying. Gellis noted that a $1.5 billion penalty may be limited in scope against companies that project much larger revenues—for example, a projection cited about $200 billion in annual revenue by 2028—though she emphasized the decision’s narrower basis on sourcing.
Fair use, competition and other court tests
Legal debate often turns on fair use factors, including purpose, amount used and market impact. Jason Henderson of JWL International said courts are inconsistent, but tend to reject training that directly competes with the original product while finding noncompeting uses more likely to qualify as fair use.
In a separate example, Judge Stephanos Bibas ruled that training on Thomson Reuters content to build a competing legal research product was not transformative and therefore not fair use in the Ross Intelligence case. That decision underscores how courts weigh the competitive effect of AI-derived services.
Authorship and unresolved questions
Other rulings address authorship of AI-generated works: in Thaler v. Perlmutter, a court held that a work entirely generated by AI is not copyrightable. Lawyers warn this raises practical problems for proving when and how much AI contributed to a work.
Most major AI firms remain involved in pending litigation, so legal standards are still evolving. Because copyright law dates from 1976 and multiple cases are unresolved, further court rulings will determine how training practices and downstream protections are ultimately defined.
Original source: TechCrunch AI