Data poisoning
Data poisoning is an adversarial-machine-learning attack in which an actor manipulates training or fine-tuning data, labels, or their selection so that the trained model behaves incorrectly at test time. Research commonly separates indiscriminate attacks that reduce overall performance, targeted attacks aimed at particular examples or classes, and backdoor attacks activated by a trigger. The term describes a broad attack family. Nightshade is one prompt-specific image-text technique within that family, not a synonym for data poisoning as a whole.
Origin and context
The literature predates modern generative AI. A broad 2022 survey traces early machine-learning poisoning work to cybersecurity research in 2006 and attacks on spam filters in 2008, then organizes later methods by attacker objective and capability. NIST's 2025 adversarial-machine-learning taxonomy places poisoning within a lifecycle-wide account of attacks and mitigations. Nightshade, introduced in 2023 and published at the 2024 IEEE Symposium on Security and Privacy, is a newer case focused on prompt-specific poisoning of text-to-image training.
Why it matters
Poisoning concerns the integrity of the learning process rather than only the inputs a deployed model receives. An evaluation can therefore show acceptable overall accuracy while missing a targeted failure learned from manipulated examples. The NIST taxonomy and independent research distinguish attacks by objectives, capabilities and lifecycle stage. Those distinctions matter when interpreting a reported result: changing labels in a controlled experiment is a different threat model from influencing a large web-collected dataset.
Example
A targeted poisoning experiment changes selected training examples so that a later model misclassifies a chosen input while retaining performance elsewhere. Nightshade studies a text-to-image variant: altered image-text training examples can create unintended associations for selected prompts under the authors' experimental conditions. This illustrates poisoning during learning. An ordinary incorrect prompt sent only to the already-trained model is not the same attack merely because its answer is wrong.
Maturity and evidence
Maturity is rated 4 for data poisoning as a research and security category. NIST's taxonomy, a multi-institution survey and University of Chicago's peer-reviewed Nightshade work use the concept across independent organizations. The rating does not imply universal attack success or a solved defense problem. Nightshade is one later case study; it neither defines the whole category nor transfers its experimental results to every generative model.
Limits and open questions
Poisoning results depend on attacker access, training-data influence, model and evaluation conditions. Success against one setup does not establish success against a different collection or training pipeline. Equally, a model error alone does not demonstrate malicious training data. Skills Intelligence separates evidence of a mechanism from claims about its prevalence or effectiveness in deployment. This entry supplies a taxonomy and a bounded case study, not attack instructions, guaranteed protection or legal advice.
Related terms
References
- Adversarial Machine Learning: A Taxonomy and Terminology of Attacks and MitigationsNational Institute of Standards and Technology · 2025-03-24 · class A
- Wild Patterns Reloaded: A Survey of Machine Learning Security against Training Data Poisoning (v3)Antonio Emanuele Cinà et al. / arXiv · 2023-03-09 · class A
- Nightshade: Prompt-Specific Poisoning Attacks on Text-to-Image Generative Models (v3; IEEE S&P 2024)Shan et al. / University of Chicago · 2024-04-29 · class A
Last updated: 2026-09-07
This term is also covered in the Skills Atlas as adversarial ai testing skill.
This term is also covered in the Skills Atlas as ai data security skill.
This term is also covered in the Skills Atlas as ai supply chain security skill.