Glossary · term

Data poisoning

Data poisoning is an adversarial-machine-learning attack in which an actor manipulates training or fine-tuning data, labels, or their selection so that the trained model behaves incorrectly at test time. Research commonly separates indiscriminate attacks that reduce overall performance, targeted attacks aimed at particular examples or classes, and backdoor attacks activated by a trigger. The term describes a broad attack family. Nightshade is one prompt-specific image-text technique within that family, not a synonym for data poisoning as a whole.

Safety2006Wave 1 · 2023Maturity: 4/5

Origin and context

The literature predates modern generative AI. A broad 2022 survey traces early machine-learning poisoning work to cybersecurity research in 2006 and attacks on spam filters in 2008, then organizes later methods by attacker objective and capability. NIST's 2025 adversarial-machine-learning taxonomy places poisoning within a lifecycle-wide account of attacks and mitigations. Nightshade, introduced in 2023 and published at the 2024 IEEE Symposium on Security and Privacy, is a newer case focused on prompt-specific poisoning of text-to-image training.

Sources: s1, s2, s3

Why it matters

Poisoning concerns the integrity of the learning process rather than only the inputs a deployed model receives. An evaluation can therefore show acceptable overall accuracy while missing a targeted failure learned from manipulated examples. The NIST taxonomy and independent research distinguish attacks by objectives, capabilities and lifecycle stage. Those distinctions matter when interpreting a reported result: changing labels in a controlled experiment is a different threat model from influencing a large web-collected dataset.

Sources: s1, s2

Example

A targeted poisoning experiment changes selected training examples so that a later model misclassifies a chosen input while retaining performance elsewhere. Nightshade studies a text-to-image variant: altered image-text training examples can create unintended associations for selected prompts under the authors' experimental conditions. This illustrates poisoning during learning. An ordinary incorrect prompt sent only to the already-trained model is not the same attack merely because its answer is wrong.

Sources: s1, s2, s3

Maturity and evidence

Maturity is rated 4 for data poisoning as a research and security category. NIST's taxonomy, a multi-institution survey and University of Chicago's peer-reviewed Nightshade work use the concept across independent organizations. The rating does not imply universal attack success or a solved defense problem. Nightshade is one later case study; it neither defines the whole category nor transfers its experimental results to every generative model.

Sources: s1, s2, s3

Limits and open questions

Poisoning results depend on attacker access, training-data influence, model and evaluation conditions. Success against one setup does not establish success against a different collection or training pipeline. Equally, a model error alone does not demonstrate malicious training data. Skills Intelligence separates evidence of a mechanism from claims about its prevalence or effectiveness in deployment. This entry supplies a taxonomy and a bounded case study, not attack instructions, guaranteed protection or legal advice.

Sources: s1, s2, s3

Related terms

References

Last updated: 2026-09-07

In the Skills Atlas

This term is also covered in the Skills Atlas as adversarial ai testing skill.

In the Skills Atlas

This term is also covered in the Skills Atlas as ai data security skill.

In the Skills Atlas

This term is also covered in the Skills Atlas as ai supply chain security skill.