AI models get convenient amnesia about source material as they grow, MIT boffins find
- ID
- 15070
- Status
- summarized
- Published
- 18 Aug 2026, 5:00 PM
- Fetched
- 18 Aug 2026, 5:34 PM
- Provider
- The Register
- Category
- technology
- Original URL
- https://www.theregister.com/ai-and-ml/2026/08/18/ai-models-get-convenient-amnesia-about-source-material-as-they-grow-mit-boffins-find/5288846
- Source URL
- https://www.theregister.com/headlines.atom
Summary
- Score
- 6.5
- Created
- 18 Aug 2026, 5:34 PM
- Tags
- Audience
- ai_ml_learnerssaas_foundersdevelopers
What happened
MIT CSAIL researchers Zheng Dai and David K Gifford found that attributing diffusion model outputs to specific training data becomes impossible as models are trained on sufficiently large corpora. Their paper, 'Outputs of Generative Diffusion Models are Often Unattributable,' to be published in Nature Communications, complicates ongoing copyright litigation like Andersen et al. v. Stability AI Ltd (2023), where plaintiffs are trying to force Midjourney to disclose training datasets.
Why it matters
If you ship products using diffusion models (image, video, audio generation), this research suggests you cannot reliably trace outputs back to specific training inputs at scale — which weakens both your ability to audit for IP infringement and plaintiffs' ability to prove it. Founders using Stable Diffusion or similar models commercially should factor this attribution gap into their copyright risk strategy rather than assuming they can retroactively identify problematic outputs.
Discussion angle
If attribution is provably impossible at scale, does this make copyright lawsuits against model providers harder to win, and does it shift liability toward the deployer who prompts the model rather than the trainer?