Technique

Knowledge Distillation

The classical Hinton-2015 formulation of distillation: minimise student-teacher logit divergence.

Definition

Originally introduced by Hinton, Vinyals & Dean in 2015, knowledge distillation trains a smaller student to match a larger teacher's soft probability distribution. It remains the theoretical foundation under which modern data-distillation variants (used for Phi, Gemma) operate.

Common use cases

  • Model compression
  • Edge inference
  • Faster serving

Related terms

    Knowledge Distillation — AI Glossary | Railwail