ToolNavs Find Useful AI Tools
Submit Sign in
Back to AI Encyclopedia
What Is SFT (Supervised Fine-Tuning)? Why Chat Models Need "Job Training" Before Going Live

What Is SFT (Supervised Fine-Tuning)? Why Chat Models Need "Job Training" Before Going Live

AI Encyclopedia • Admin • • 6 views

SFT (Supervised Fine-Tuning) is the "job training" a chat model goes through before going live: using thousands of "instruction-response" examples, it teaches a pretrained large model to speak and act the way humans expect. Note that it does not pump new knowledge into the model; it teaches the model to "say what it already knows, in the right format."

What pretraining, SFT, and RLHF each do

A chat model's training pipeline usually runs in three stages: pretraining → SFT → RLHF (or alignment methods like DPO), each with a completely different job:

StageWhat it learnsIn one line
PretrainingLanguage patterns and world knowledge from massive textLearns to "talk and understand"
SFTFrom instruction examples: follow instructions and answer in the expected formatLearns to "do as asked"
RLHF / DPOFrom human preferences: be safer and more helpfulLearns to "please the user"

How SFT is done

First, prepare an instruction dataset: thousands of "user instruction – ideal response" pairs covering Q&A, writing, coding, reasoning, and more. Second, the loss is computed only on the response: the model sees the whole example but is only graded on how closely its answer matches the demonstration — the instruction part doesn't count. Third, quality beats quantity by far: research has shown that a few hundred to a thousand high-quality demonstrations often outperform hundreds of thousands of sloppy ones.

Incidentally, SFT doesn't always require full fine-tuning. Parameter-efficient methods like LoRA fine-tuning can do SFT at low cost, and are the common choice for individuals and small teams.

Where SFT hits its ceiling

First, it can't learn knowledge newer than its training cutoff: the knowledge ceiling is set by pretraining; SFT only teaches "how to express it." Second, the behavioral-cloning ceiling: the model can at most match the quality of its demonstration data — it can't outperform its teachers. Third, biases in the data get amplified: the biases and mistakes in the demonstration data get learned and reinforced by the model.

Three common misconceptions

Misconception 1: "More SFT data is always better." Wrong. Bad data teaches bad habits; high-quality demonstrations are what matter.

Misconception 2: "The model gets smarter after SFT." Wrong. The model's knowledge doesn't change; it just gets better at answering. The capability ceiling was already set during pretraining.

Misconception 3: "With SFT, you no longer need prompt engineering or RAG." Wrong. SFT sets the model's "default behavior," RAG handles "real-time knowledge," and prompt engineering handles "the task at hand" — three layers that each mind their own business and can't replace one another.

In essence, SFT is a "professional-habits training" for the model: it doesn't change what the model knows, only how it answers.

Recommended Tools

More