Glossary · term

Distillation attack

A distillation attack is the adversarial use of outputs from a service-accessed teacher model to train a separate student that reproduces selected behavior or capabilities, usually through systematic black-box queries. The attack label describes the acquisition context, not knowledge distillation itself. Current usage overlaps with model extraction, but it does not require copying weights and should not be applied merely because one model learns from another with permission.

Safety2026-02-12Wave 3 · 2025–26Maturity: 3/5

Origin and context

Model extraction was established as a research problem by Tramèr and colleagues in 2016. An ACL 2025 paper then treated distillation as a method for extracting LLM behavior. The earliest exact distillation-attack usage verified here is Google's February 2026 threat report; Anthropic used it independently later that month, and Helen Toner's April Senate testimony carried it into policy discussion. This sequence supports adoption, not a coinage claim. Incident attributions and volumes in provider reports remain those providers' findings.

Sources: s1, s2, s3, s4, s5

Why it matters

A public model API exposes a behavioral interface even when weights and original training data remain private. Large, targeted query sets can become synthetic training data for a student, reducing some data-generation and experimentation costs. Providers may lose differentiated capabilities or control over safety restrictions, while defenders still have to infer intent from traffic. Recent research therefore models attacker query budget, data budget and interface profile instead of treating every high-volume or training-related request as hostile.

Sources: s1, s2, s6

Example

If a lab distils its own teacher into a smaller deployment model, or trains from another model's outputs under an applicable permission, that is ordinary distillation. A contrasting pattern is coordinated accounts or proxies sending repeated, capability-focused queries and aggregating the responses to train a competing student while evading access restrictions. Anthropic describes that pattern, while Google describes related extraction through legitimate API access. The mechanics alone do not prove copied weights, copyright infringement, trade-secret misappropriation or any named actor's liability; those are separate factual and legal questions.

Sources: s1, s2, s5, s7

How it differs

Knowledge distillation

Knowledge distillation is a teacher-student training method and can be routine, authorized engineering. A distillation attack is a threat-model label for adversarial acquisition using that method. Permission, deception, access circumvention, scale and extraction purpose may inform the label, but they are not properties of the optimization method itself.

Maturity and evidence

Maturity is rated 3. The exact label appears independently in Google and Anthropic operational reports, in policy testimony and in 2026 research, while its technical core inherits a decade of model-extraction work. It is not rated 4 because Google largely equates it with model extraction, Anthropic emphasizes coordinated evasive behavior, and emerging defense research still lacks a shared threat model.

Sources: s1, s2, s3, s5, s6

Limits and open questions

Provider incident reports are first-party accounts and should not be converted into findings by a court or regulator. Their terms-of-service claims apply to their own services; the label itself does not settle copyright, trade-secret, contract, computer-misuse or competition questions. The U.S. Copyright Office likewise treats AI-training analysis as specific to the use and circumstances, rather than a bright-line answer. Detection can also flag authorized research or large synthetic-data jobs. This entry is technical context, not legal advice.

Sources: s1, s2, s5, s6, s7

Related terms

References

Last updated: 2026-09-05

In the Skills Atlas

This term is also covered in the Skills Atlas as knowledge distillation skill.

In the Skills Atlas

This term is also covered in the Skills Atlas as ai risk management skill.