- Type
- Training failure mode in generative models
- Named by
- Shumailov et al., Oxford, Cambridge and Toronto researchers
- Key publication
- Nature, July 2024
- Cause
- Recursive training on model-generated data
- Main mitigation
- Retaining and accumulating human-originated data
- Related
- Synthetic Data, Generative AI, Data Pipeline
- Type
- Training failure mode in generative models
- Named by
- Shumailov et al., Oxford, Cambridge and Toronto researchers
- Key publication
- Nature, July 2024
- Cause
- Recursive training on model-generated data
- Main mitigation
- Retaining and accumulating human-originated data
- Related
- Synthetic Data, Generative AI, Data Pipeline
History
The concept emerged as synthetic data spread through machine learning pipelines. In "The Curse of Recursion" (2023), researchers showed that language models trained on text generated by other models forget rare events and converge toward repetitive output. The 2024 Nature paper extended the demonstration across Gaussian mixture models, variational autoencoders and large language models, finding that even training runs that begin well degrade as data are recycled, with the tails of the original distribution disappearing first.[1][2] A widely cited follow-up by Gerstgrasser and colleagues (2024) refined the picture: collapse is driven by replacing real data with synthetic data, while accumulating successive generations alongside the original data avoids it, suggesting the effect is a data-management problem as much as a mathematical inevitability.[3] The debate intensified as web-scale training corpora became increasingly contaminated by AI-generated text and images, and as researchers warned about a finite supply of high-quality human data.[4]
Key Concepts and Technology
Model collapse arises from compounding errors across training generations. Each round of generation and re-training introduces three distortions: statistical sampling error from finite generated samples, expressivity limits that cap what a model family can represent, and approximation error from imperfect fitting. Rare but real examples vanish first, so successive models learn a narrowed, distorted picture of the world. In language models the symptoms include repetitive phrasing, lost cultural and factual detail, and rising perplexity on original data. Mitigations follow directly from the diagnosis: keep human-originated data in every training mix, accumulate rather than replace historical data, label and track synthetic content through data provenance standards, apply watermarks such as SynthID, and invest in curation and filtering. Researchers also note that limited, carefully controlled use of synthetic data — for example in distillation or data augmentation — can still improve models when anchored to real data.[1][3]
Applications and Impact
Model collapse shapes how AI laboratories source data: it is part of the economic logic behind licensing deals for human content, investment in data curation pipelines, and caution about training on scraped web data of unknown provenance. For organisations adopting generative AI, the concept argues for disciplined data governance — recording where training data came from, preserving authentic records, and monitoring for quality decay across model versions. For the public, it reframes today's web as a shared training commons that risks being exhausted or corrupted if AI-generated content crowds out human material.[2][4]
>See Also
For Malaysia, model collapse sharpens the value of authentic local data. Malay, Tamil, Mandarin and indigenous-language corpora are small relative to English, so recycled synthetic content could degrade the quality of local-language AI faster than in high-resource settings; national language-model efforts rely on curated public and licensed text to avoid this trap. Government agencies and enterprises generating large volumes of AI-drafted content should document what is human-originated, and research institutions can contribute by building verified Malaysian datasets. Under the Personal Data Protection Act 2010, organisations also need lawful handling of personal data used in training, reinforcing the case for the provenance tools that national AI governance initiatives such as the Malaysia AI Governance Framework already promote.[3][4]
References
- ↑Shumailov, I., et al. (2024). AI models collapse when trained on recursively generated data. Nature 631, 755-759. https://www.nature.com/articles/s41586-024-07566-y
- ↑Shumailov, I., et al. (2023). The Curse of Recursion: Training on Generated Data Makes Models Forget. arXiv:2305.17493. https://arxiv.org/abs/2305.17493
- ↑Gerstgrasser, M., et al. (2024). Is Model Collapse Inevitable? Breaking the Curse of Recursion by Accumulating Real and Synthetic Data. arXiv:2404.01413. https://arxiv.org/abs/2404.01413
- ↑Nature News. (2024). AI models fed AI-generated data quickly spew nonsense. https://www.nature.com/articles/d41586-024-02355-z