01
Factual background & dispute
- Authors alleged that Anthropic downloaded extensive volumes of books from repositories including LibGen and PiLiMi, using portions of these works for model training.
- In its partial summary judgment ruling, the court separately analyzed model training, internal digitization of purchased print books, and long-term retention of works obtained from pirated repositories.
- The parties reached a class action settlement covering designated works, focusing on past input and copying conduct.
02
Core issues & judicial focus
- Whether training large language models on copyrighted books constitutes fair use
- How the conversion of purchased print books into internal digital copies should be evaluated
- Whether downloading and retaining books from unauthorized repositories creates independent liability
03
Judicial finding & holding
- The court held that model training and the digitization of lawfully purchased print books constituted fair use on the record presented.
- The court held that obtaining and storing books from pirated repositories required independent scrutiny that downstream training purposes could not automatically excuse.
- In July 2026, the court granted final approval to a $1.5 billion settlement covering 482,460 works listed in the settlement registry.
04
Practical risk implications
01Fair use assessments for AI training must examine data acquisition methods alongside subsequent data retention practices.
02Enterprises should maintain comprehensive records of purchases, licenses, downloads, deduplication, deletions, and training runs.
03Settlement agreements resolve only defined past input activities; downstream output disputes and future conduct require independent assessment.