AIWiki
Malaysia
Back to all articles
AI Foundationsmodel collapsesynthetic datadata quality

Model Collapse

4 min readUpdated September 2026
Model Collapse
Type
Training failure mode in generative models
Named by
Shumailov et al., Oxford, Cambridge and Toronto researchers
Key publication
Nature, July 2024
Cause
Recursive training on model-generated data
Main mitigation
Retaining and accumulating human-originated data
Related
Synthetic Data, Generative AI, Data Pipeline
Model collapse is a degenerative failure mode in generative artificial intelligence where models trained on data produced by earlier generations of models progressively lose accuracy and diversity, drifting away from the real-world distribution they were meant to learn. The phenomenon was named and demonstrated at scale in a 2024 paper in Nature by Ilia Shumailov and colleagues, following an earlier preprint that introduced it as the "curse of recursion".[1][2]

History

The concept emerged as synthetic data spread through machine learning pipelines. In "The Curse of Recursion" (2023), researchers showed that language models trained on text generated by other models forget rare events and converge toward repetitive output. The 2024 Nature paper extended the demonstration across Gaussian mixture models, variational autoencoders and large language models, finding that even training runs that begin well degrade as data are recycled, with the tails of the original distribution disappearing first.[1][2] A widely cited follow-up by Gerstgrasser and colleagues (2024) refined the picture: collapse is driven by replacing real data with synthetic data, while accumulating successive generations alongside the original data avoids it, suggesting the effect is a data-management problem as much as a mathematical inevitability.[3] The debate intensified as web-scale training corpora became increasingly contaminated by AI-generated text and images, and as researchers warned about a finite supply of high-quality human data.[4]

Key Concepts and Technology

Model collapse arises from compounding errors across training generations. Each round of generation and re-training introduces three distortions: statistical sampling error from finite generated samples, expressivity limits that cap what a model family can represent, and approximation error from imperfect fitting. Rare but real examples vanish first, so successive models learn a narrowed, distorted picture of the world. In language models the symptoms include repetitive phrasing, lost cultural and factual detail, and rising perplexity on original data. Mitigations follow directly from the diagnosis: keep human-originated data in every training mix, accumulate rather than replace historical data, label and track synthetic content through data provenance standards, apply watermarks such as SynthID, and invest in curation and filtering. Researchers also note that limited, carefully controlled use of synthetic data — for example in distillation or data augmentation — can still improve models when anchored to real data.[1][3]

Applications and Impact

Model collapse shapes how AI laboratories source data: it is part of the economic logic behind licensing deals for human content, investment in data curation pipelines, and caution about training on scraped web data of unknown provenance. For organisations adopting generative AI, the concept argues for disciplined data governance — recording where training data came from, preserving authentic records, and monitoring for quality decay across model versions. For the public, it reframes today's web as a shared training commons that risks being exhausted or corrupted if AI-generated content crowds out human material.[2][4]

>See Also

🇲🇾Malaysian Context

For Malaysia, model collapse sharpens the value of authentic local data. Malay, Tamil, Mandarin and indigenous-language corpora are small relative to English, so recycled synthetic content could degrade the quality of local-language AI faster than in high-resource settings; national language-model efforts rely on curated public and licensed text to avoid this trap. Government agencies and enterprises generating large volumes of AI-drafted content should document what is human-originated, and research institutions can contribute by building verified Malaysian datasets. Under the Personal Data Protection Act 2010, organisations also need lawful handling of personal data used in training, reinforcing the case for the provenance tools that national AI governance initiatives such as the Malaysia AI Governance Framework already promote.[3][4]

References

  1. ↑Shumailov, I., et al. (2024). AI models collapse when trained on recursively generated data. Nature 631, 755-759. https://www.nature.com/articles/s41586-024-07566-y
  2. ↑Shumailov, I., et al. (2023). The Curse of Recursion: Training on Generated Data Makes Models Forget. arXiv:2305.17493. https://arxiv.org/abs/2305.17493
  3. ↑Gerstgrasser, M., et al. (2024). Is Model Collapse Inevitable? Breaking the Curse of Recursion by Accumulating Real and Synthetic Data. arXiv:2404.01413. https://arxiv.org/abs/2404.01413
  4. ↑Nature News. (2024). AI models fed AI-generated data quickly spew nonsense. https://www.nature.com/articles/d41586-024-02355-z