Knowledge distillation
Knowledge distillation is a training method in which a student model learns from the outputs or internal representations of a teacher model, often alongside ground-truth labels. Soft probability targets can convey relationships between classes that hard labels omit. The goal is usually to transfer useful behavior into a smaller or more efficient student. It is distinct from ordinary quantization, pruning, or copying model weights.
Origin and context
Hinton, Vinyals, and Dean introduced the durable modern formulation in a paper submitted in March 2015, showing how an ensemble's knowledge could be compressed into a single model using softened outputs. DistilBERT later adapted the idea to pretrained language models with a triple loss combining language modeling, distillation, and representation alignment. Gou and colleagues' survey, first posted in 2020 and accepted by the International Journal of Computer Vision in 2021, organized teacher-student architectures, knowledge forms, algorithms, applications, and open challenges.
Why it matters
Distillation separates the model used to supply training supervision from the model deployed for a task. A smaller student can reduce inference demands, but the achievable trade-off depends on the student architecture, the teacher signal and the training data. The survey also describes uses beyond compression. For an implementation decision, compare the student with its teacher and a student trained without distillation; otherwise a smaller model alone does not demonstrate the contribution of the method.
Example
A team can run a large classifier over a curated corpus, save its probability distribution for each example, and train a smaller student on both the original labels and those soft targets. DistilBERT is a language-model example: its authors reported a model 40 percent smaller that retained 97 percent of measured language-understanding performance and ran 60 percent faster in their evaluated setup. Those figures are study-specific, not universal expectations.
Maturity and evidence
Maturity is rated 4. The original Google work and Hugging Face's DistilBERT provide distinct organizational examples, while the independent survey maps applications across architectures and learning settings. This supports an established method, not a universal compression result. The rating describes technical adoption and does not confer regulatory status or assurance about a particular distilled model.
Limits and open questions
Distillation is defined by a learning objective, not by an assertion about permission to use a teacher or its outputs. The scientific method alone cannot settle those separate questions. Technically, choosing which outputs or representations to match remains important: the survey identifies open questions about teacher-student architecture and generalization. A student's benchmark result should not be treated as evidence that it reproduces every teacher capability.
Related terms
References
- Distilling the Knowledge in a Neural NetworkGoogle / arXiv · 2015-03-09 · class A
- DistilBERT, a distilled version of BERT: smaller, faster, cheaper and lighterHugging Face / arXiv · 2019-10-02 · class A
- Knowledge Distillation: A Survey (version 7; accepted by IJCV)Gou, Yu, Maybank and Tao / arXiv · 2021-05-20 · class A
Last updated: 2026-09-05
This term is also covered in the Skills Atlas as knowledge distillation skill.
This term is also covered in the Skills Atlas as model pruning skill.
This term is also covered in the Skills Atlas as model quantization skill.