Technique
Knowledge Distillation
The classical Hinton-2015 formulation of distillation: minimise student-teacher logit divergence.
Definition
Originally introduced by Hinton, Vinyals & Dean in 2015, knowledge distillation trains a smaller student to match a larger teacher's soft probability distribution. It remains the theoretical foundation under which modern data-distillation variants (used for Phi, Gemma) operate.
Common use cases
- Model compression
- Edge inference
- Faster serving