AI Weekly Malaysia

Back to items Summaries

Is it legal to train AI models on copyrighted books? It’s complicated

ID
17051
Status
summarized
Published
23 Aug 2026, 11:00 PM
Fetched
23 Aug 2026, 11:19 PM
Provider
TechCrunch
Category
technology
Original URL
https://techcrunch.com/2026/08/23/is-it-legal-to-train-ai-models-on-copyrighted-books-its-complicated/
Source URL
https://techcrunch.com/feed/

Summary

Score
6.5
Created
23 Aug 2026, 11:20 PM
Tags
Audience
developersai_ml_learnerssaas_founders

What happened

Judge William Alsup ruled that Anthropic's AI training on copyrighted books was lawful, comparing LLM ingestion to a writer studying literature, but still ordered a $1.5 billion penalty because Anthropic pirated the books from illegal shadow libraries. Attorney Cathy Gellis notes the ruling is effectively good news for AI companies, since the fine is small relative to projected revenue and the core legal question—whether training constitutes copying—was answered in AI companies' favor.

Why it matters

If you're building AI products or training models, the legal risk now hinges on HOW you acquire training data, not whether you train on copyrighted works. The ruling suggests scraping legally accessible content for training may be defensible, but downloading from pirate sources carries massive financial liability. Founders should audit their data pipelines for provenance before scaling, and SaaS builders using third-party models should check whether their providers' training data sourcing could expose them to downstream risk.

Discussion angle

The ruling draws a line between 'reading' copyrighted works (legal training) and 'copying' them (illegal piracy)—but does this analogy hold up for smaller builders who can't absorb a $1.5B fine, and what does it mean for open-source model trainers who rely on datasets of unclear origin?

Top