Big models learn well but run slowly. Small models run fast but learn less. Knowledge distillation splits the difference by training a compact student model to imitate a larger teacher, using the teacher's soft predictions rather than only hard labels.
Soft predictions carry more information. When a teacher assigns 70 percent probability to cat and 25 percent to dog, the student learns that cat and dog share features. Hard labels would say only cat. That extra signal helps the student generalize better than training on labels alone.
Common distillation variants
- Response-based distillation from output logits
- Feature-based distillation from intermediate layers
- Relation-based distillation from layer relationships
- Self-distillation within the same architecture
The technique powers many production systems. Mobile assistants, recommendation models, and real-time detectors often rely on distilled students that retain most of the teacher's accuracy at a fraction of the cost.
Comments (2)
Leave a comment