Is it legal to train AI models on copyrighted books? It’s complicated
- ID
- 17051
- Status
- summarized
- Published
- 23 Aug 2026, 11:00 PM
- Fetched
- 23 Aug 2026, 11:19 PM
- Provider
- TechCrunch
- Category
- technology
- Original URL
- https://techcrunch.com/2026/08/23/is-it-legal-to-train-ai-models-on-copyrighted-books-its-complicated/
- Source URL
- https://techcrunch.com/feed/
Summary
- Score
- 6.5
- Created
- 23 Aug 2026, 11:20 PM
- Tags
- Audience
- developersai_ml_learnerssaas_founders
What happened
Judge William Alsup ruled that Anthropic's AI training on copyrighted books was lawful, comparing LLM ingestion to a writer studying literature, but still ordered a $1.5 billion penalty because Anthropic pirated the books from illegal shadow libraries. Attorney Cathy Gellis notes the ruling is effectively good news for AI companies, since the fine is small relative to projected revenue and the core legal question—whether training constitutes copying—was answered in AI companies' favor.
Why it matters
If you're building AI products or training models, the legal risk now hinges on HOW you acquire training data, not whether you train on copyrighted works. The ruling suggests scraping legally accessible content for training may be defensible, but downloading from pirate sources carries massive financial liability. Founders should audit their data pipelines for provenance before scaling, and SaaS builders using third-party models should check whether their providers' training data sourcing could expose them to downstream risk.
Discussion angle
The ruling draws a line between 'reading' copyrighted works (legal training) and 'copying' them (illegal piracy)—but does this analogy hold up for smaller builders who can't absorb a $1.5B fine, and what does it mean for open-source model trainers who rely on datasets of unclear origin?