Pre-training is the first stage of training a large language model: before it learns to chat or write code, the model does one thing over massive amounts of text — it repeatedly guesses the next word. This stage does not teach it politeness or instruction-following; it only bakes the statistical patterns of language, factual associations, and expression habits into its parameters. The conversational ability and reasoning style you are familiar with all stand on the foundation that pre-training lays down.
The training task is only one: guess the next word
The task is surprisingly plain. Give the model a passage, mask the word that follows, and let it guess; if it is wrong, the parameters are adjusted, and if it is right, they are reinforced. After trillions of such exercises, the model gradually learns the connections between grammar, common sense, and technical terminology.
This approach is called self-supervised learning: the answers are hidden inside the text itself, so no human labeling is needed item by item. That is also why pre-training can scale so far — web pages, books, papers, and code repositories on the internet all become ready-made textbooks once cleaned.
What it reads matters more than how much
The quality of the textbook sets the ceiling. Labs follow broadly similar recipes: gather raw corpora from public web pages, books, encyclopedias, and code, then deduplicate, filter low-quality content, and strip private information. The clear trend in recent years is "less but better": rather than feeding in the whole noisy web, teams raise the share of textbook-like and code data.
So when you evaluate a model, "how many tokens it trained on" is only a scale metric, not a direct measure of quality. Two models with similar parameter counts but different data mixes can differ by a full tier in coding and math performance.
Pre-training versus post-training
The result of pre-training is called a base model. It knows a lot, but behaves like a "continuation machine": ask it a question and it may continue by inventing a Q&A passage instead of answering you. Turning it into today's responsive assistant that can decline inappropriate requests takes two further steps: first supervised fine-tuning with human demonstrations, then alignment with human preference feedback.
A handy way to remember the split: pre-training decides how much it knows and how solid the foundation is; post-training decides how it deals with people. The same base model can become assistants with very different styles under different post-training recipes.
Three common misunderstandings
First, pre-training is not memorizing original texts into a database. The model stores parameterized statistical patterns, not a searchable text library; it can recite famous passages fluently because they appear so often in the corpus, not because it saved the files. This is also one reason it can be confidently wrong: it generates plausible-sounding statements, not retrieved facts.
Second, finishing pre-training does not mean training is over. Using a base model directly as a chat product usually feels poor; when vendors announce a "model", the full post-training pipeline is included by default.
Third, pre-training is not a one-time education for life. Major labs keep training on new data and releasing new versions; when a model feels smarter, much of the credit often goes to a new round of pre-training data and its mix, not just parameter tweaks.