Glossary · term

Model collapse

Model collapse is a degenerative process in which generative models trained recursively on model-produced data lose information about the original data distribution. Early effects can erase low-probability events and reduce diversity; later effects can make the learned distribution converge toward a distorted, low-variance approximation. It is a training-data feedback problem, not a claim that every use of synthetic data inevitably ruins a model.

Training2023-05-31Wave 1 · 2023Maturity: 3/5

Origin and context

Shumailov and colleagues used Model Collapse in the May 31, 2023 revision of their preprint The Curse of Recursion. Its first submission, four days earlier, had used different terminology. Their research later appeared in Nature in July 2024. Separately, Alemohammad and colleagues introduced Model Autophagy Disorder (MAD) in a July 2023 preprint about self-consuming generative-image training loops. The terms describe overlapping failure phenomena, but the papers investigate different experimental and analytical settings.

Sources: s1, s2, s3

Why it matters

A generated training example is a sample from a learned approximation, not a fresh observation of the original distribution. When later generations increasingly learn from such samples, mistakes in estimating rare events can feed back into the next model. The practical question is therefore not simply whether a dataset contains synthetic material. It is whether each generation replaces, retains, or supplements earlier observations, and whether evaluation detects losses in diversity as well as average quality.

Sources: s1, s3, s4

Example

As a Skills Intelligence illustration, compare two image-training pipelines. The first discards its original photographs and retrains each generation only on the previous generator's output. The second retains the original photographs and accumulates additional synthetic examples. These are not equivalent recursive loops. Gerstgrasser and colleagues' 2024 preprint reported collapse in replacement settings but avoided it in the accumulation settings they tested, including language, image, and molecular data. That result supports a conditional comparison, not a guarantee for every accumulated dataset.

Sources: s3, s4

Maturity and evidence

Skills Intelligence rates the concept at maturity 3: it has an established research definition, peer-reviewed evidence, and independent experiments examining when it does and does not arise. These sources establish a research phenomenon rather than broad deployment of a standard prevention method. The rating therefore does not infer operational maturity from citation visibility, or turn a finding under particular assumptions into a prediction that future AI models must deteriorate.

Sources: s1, s2, s3, s4

Limits and open questions

Not every quality decline is model collapse: a faulty fine-tuning run, distribution shift, or low-quality source data may have another explanation. The MAD experiments emphasize access to fresh real data, whereas accumulation research shows that retaining original data can change the outcome under its tested conditions. Model family, sampling, dataset replacement, and evaluation all affect the result. Neither study establishes a universal safe proportion of synthetic training data.

Sources: s3, s4

Related terms

References

Last updated: 2026-09-05

In the Skills Atlas

This term is also covered in the Skills Atlas as training data curation skill.

In the Skills Atlas

This term is also covered in the Skills Atlas as data quality management skill.

In the Skills Atlas

This term is also covered in the Skills Atlas as synthetic data generation skill.