ToolNavs Find Useful AI Tools
Submit Sign in
Back to AI Encyclopedia
Knowledge Distillation: How Big Models Teach Small Ones — Principles, Limits, and Three Myths

Knowledge Distillation: How Big Models Teach Small Ones — Principles, Limits, and Three Myths

AI Encyclopedia • Admin • • 2 views

Knowledge distillation is a technique that lets a small model (the student) learn from an already-trained large model (the teacher). It was formalized by Geoffrey Hinton and colleagues in their 2015 paper "Distilling the Knowledge in a Neural Network": the goal is to let a smaller, faster model keep as much of the large model's ability as possible.

The core mechanism: the student learns the teacher's "hesitation," not just the answer

Ordinary training uses "hard labels": this image is a cat, so record it as 100% cat. The teacher gives "soft labels" instead: 70% cat, 20% tiger, 10% something else. Hinton called this information "dark knowledge" — the similarity structure between classes is what the teacher actually learned.

To expose these faint probability signals, distillation introduces a temperature parameter T: divide the teacher's logits by T before applying softmax. The larger T is, the "softer" the distribution becomes, and secondary classes that were squashed to 0.1% get lifted so the student can see the teacher's decision structure. The student's training loss is a weighted mix of two terms: a KL divergence that follows the teacher's soft labels, and a standard cross-entropy loss against hard labels.

The key ingredients of knowledge distillation

A complete distillation run needs four things: a trained teacher model (usually a large model), a student model to train (a small architecture), the temperature parameter T, and the weighting between the soft and hard loss terms. The teacher "writes the exam and provides the reference answers," while the student finds the balance between mimicking the teacher's output distribution and staying faithful to the true labels.

The ceiling of distillation: when it's worse than training directly

Distillation is not a magic wand. First, the student's ceiling is roughly set by the teacher: if the teacher is weak on some task, the student only learns to "be wrong like the teacher." Second, a student with too little capacity "can't learn" — the teacher's knowledge won't fit into a tiny network. Third, distillation has an upfront cost: generating soft labels means running the big teacher model first. Finally, if the target task lies outside the teacher's distribution, it's better to train the small model directly on the target data.

Three common misconceptions

Myth 1: distillation is compression, the same as quantization or pruning. All three make models smaller and faster, but distillation happens during training and transfers capability; quantization happens at deployment and lowers numerical precision; pruning deletes redundant parameters. They can be stacked (distill first, then quantize is a common combo), but they can't replace each other.

Myth 2: the student can never beat the teacher. Usually the teacher is indeed the ceiling, but the loss also mixes in true hard labels, so after fine-tuning on task data, the student can occasionally surpass the teacher on a single metric.

Myth 3: distillation is lossless. Almost every distillation loses a little performance; the question is how much and where. The flip side of "retains 97% of performance" is that the missing 3% tends to concentrate in long-tail, hard examples.

Where distillation, quantization, and pruning actually differ

In one sentence: distillation changes "how the model learns" and produces a brand-new small model; quantization changes "how numbers are stored" while the architecture stays; pruning changes "how much structure survives" by deleting parameters directly. They act at different stages of a model's life cycle, often appear together, but solve different problems. For a more accessible take, see our earlier piece Model Distillation: Why More "Small Models" Are Catching Up to Big Ones.

Real-world applications

The classic example is HuggingFace's DistilBERT: with BERT as teacher, the student model has 40% fewer parameters, runs 60% faster, and keeps about 97% of the language understanding ability. The LLM era has its own showcase: the DeepSeek-R1 team distilled DeepSeek-R1-Distill-Qwen-7B from 800K reasoning samples generated by R1, scoring 55.5% pass rate on the AIME 2024 math test — far above directly trained models of the same size. Today, the small models running on-device voice assistants, local customer service, and in-car systems mostly have distillation behind them.

Recommended Tools

More