SFT (Supervised Fine-Tuning) is the "job training" a chat model goes through before going live: using thousands of "instruction-response" examples, it teaches a pretrained large model to speak and act the way humans expect. Note that it does not pump new knowledge into the model; it teaches the model to "say what it already knows, in the right format."
What pretraining, SFT, and RLHF each do
A chat model's training pipeline usually runs in three stages: pretraining → SFT → RLHF (or alignment methods like DPO), each with a completely different job:
| Stage | What it learns | In one line |
|---|---|---|
| Pretraining | Language patterns and world knowledge from massive text | Learns to "talk and understand" |
| SFT | From instruction examples: follow instructions and answer in the expected format | Learns to "do as asked" |
| RLHF / DPO | From human preferences: be safer and more helpful | Learns to "please the user" |
How SFT is done
First, prepare an instruction dataset: thousands of "user instruction – ideal response" pairs covering Q&A, writing, coding, reasoning, and more. Second, the loss is computed only on the response: the model sees the whole example but is only graded on how closely its answer matches the demonstration — the instruction part doesn't count. Third, quality beats quantity by far: research has shown that a few hundred to a thousand high-quality demonstrations often outperform hundreds of thousands of sloppy ones.
Incidentally, SFT doesn't always require full fine-tuning. Parameter-efficient methods like LoRA fine-tuning can do SFT at low cost, and are the common choice for individuals and small teams.
Where SFT hits its ceiling
First, it can't learn knowledge newer than its training cutoff: the knowledge ceiling is set by pretraining; SFT only teaches "how to express it." Second, the behavioral-cloning ceiling: the model can at most match the quality of its demonstration data — it can't outperform its teachers. Third, biases in the data get amplified: the biases and mistakes in the demonstration data get learned and reinforced by the model.
Three common misconceptions
Misconception 1: "More SFT data is always better." Wrong. Bad data teaches bad habits; high-quality demonstrations are what matter.
Misconception 2: "The model gets smarter after SFT." Wrong. The model's knowledge doesn't change; it just gets better at answering. The capability ceiling was already set during pretraining.
Misconception 3: "With SFT, you no longer need prompt engineering or RAG." Wrong. SFT sets the model's "default behavior," RAG handles "real-time knowledge," and prompt engineering handles "the task at hand" — three layers that each mind their own business and can't replace one another.
In essence, SFT is a "professional-habits training" for the model: it doesn't change what the model knows, only how it answers.