Glossary · term

Synthetic data

Synthetic data is artificially generated information designed to reproduce selected properties or support tasks normally served by observed data. It may come from statistical models, simulators, rules, or generative AI. In machine learning it can augment, replace, rebalance, or label parts of a training set. Synthetic does not mean anonymous, unbiased, accurate, or safe by default; those properties require separate evidence.

Training1993Wave 1 · 2023Maturity: 4/5

Origin and context

Drechsler and Haensch's 2023 review preprint traces the modern proposal to Donald Rubin's 1993 work on producing synthetic records for disclosure limitation. The concept later expanded across privacy engineering and machine learning. In 2023, Self-Instruct demonstrated a prominent language-model workflow: a model generated instruction, input, and output examples, filtered them, and was fine-tuned on the resulting data. That application is influential but is not the origin of synthetic data as a category.

Sources: s1, s2

Why it matters

Synthetic data can create rare scenarios, rebalance classes, lower collection costs, and support experimentation when direct access to sensitive or scarce observations is constrained. It can also inherit bias, leak information about source records, introduce factual errors, or create feedback loops when later models repeatedly consume earlier model outputs. NIST warns that synthetic data without differential privacy does not provide a robust privacy guarantee and may reduce accuracy for subgroups.

Sources: s1, s2, s3, s4

Example

As an illustrative workflow, a support team can generate candidate conversations, filter invalid or near-duplicate examples, and compare a model trained with them against a baseline on separately collected test cases. Human checking does not make those generated examples observed data. This is an application of the generation-and-filtering pattern, not a reported experiment or a privacy guarantee. Recursive training on model outputs is a further design choice, not a necessary feature of every synthetic dataset.

Sources: s2, s3, s4

How it differs

Synthetic data flywheel

Synthetic-data flywheel describes an iterative workflow around generation, selection and later training; synthetic data names the generated material. A dataset can be synthetic without participating in a loop. Repetition alone does not establish a beneficial flywheel: filtering and evaluation determine whether later training data are useful, and recursive self-consumption can degrade a model under the conditions studied in model-collapse research.

Maturity and evidence

Maturity is rated 4. The historical review documents operational synthetic-data projects at organizations including the U.S. Census Bureau and Statistics New Zealand; Self-Instruct supplies a different, language-model application. This is evidence of adoption beyond one research team. It does not make all generators equivalent or turn the label into a privacy certification. NIST guidance separates the privacy mechanism from the synthetic appearance of records.

Sources: s1, s2, s3

Limits and open questions

A synthetic dataset can preserve a useful pattern while losing other relationships or reducing accuracy for subgroups. NIST explains that generation introduces additional uncertainty and that non-differentially-private synthesis may remain vulnerable to privacy attacks. Model-collapse results concern recursive training conditions; they do not show that every use of generated examples must fail. For evaluation, specify what properties need to be retained and test those properties independently of how realistic individual examples look.

Sources: s1, s3, s4

Related terms

References

Last updated: 2026-09-05

In the Skills Atlas

This term is also covered in the Skills Atlas as synthetic data generation skill.

In the Skills Atlas

This term is also covered in the Skills Atlas as data augmentation skill.

In the Skills Atlas

This term is also covered in the Skills Atlas as training data curation skill.