AI Weekly Malaysia

Back to items Summaries

AI models get convenient amnesia about source material as they grow, MIT boffins find

ID
15070
Status
summarized
Published
18 Aug 2026, 5:00 PM
Fetched
18 Aug 2026, 5:34 PM
Provider
The Register
Category
technology
Original URL
https://www.theregister.com/ai-and-ml/2026/08/18/ai-models-get-convenient-amnesia-about-source-material-as-they-grow-mit-boffins-find/5288846
Source URL
https://www.theregister.com/headlines.atom

Summary

Score
6.5
Created
18 Aug 2026, 5:34 PM
Tags
Audience
ai_ml_learnerssaas_foundersdevelopers

What happened

MIT CSAIL researchers Zheng Dai and David K Gifford found that attributing diffusion model outputs to specific training data becomes impossible as models are trained on sufficiently large corpora. Their paper, 'Outputs of Generative Diffusion Models are Often Unattributable,' to be published in Nature Communications, complicates ongoing copyright litigation like Andersen et al. v. Stability AI Ltd (2023), where plaintiffs are trying to force Midjourney to disclose training datasets.

Why it matters

If you ship products using diffusion models (image, video, audio generation), this research suggests you cannot reliably trace outputs back to specific training inputs at scale — which weakens both your ability to audit for IP infringement and plaintiffs' ability to prove it. Founders using Stable Diffusion or similar models commercially should factor this attribution gap into their copyright risk strategy rather than assuming they can retroactively identify problematic outputs.

Discussion angle

If attribution is provably impossible at scale, does this make copyright lawsuits against model providers harder to win, and does it shift liability toward the deployer who prompts the model rather than the trainer?

Top