RAG (Retrieval-Augmented Generation) is a technique that lets a large language model look things up before answering: before generating a response, the model retrieves content relevant to the question from an external knowledge base, then writes its answer based on that material. It targets two chronic weaknesses of LLMs — training data with a cutoff date, and the tendency to confidently fabricate what they don't know (hallucinations) — shifting the model from "answering from memory" to "answering while looking things up."
What Are the Three Steps?
The full pipeline has just three steps, all named in the acronym:
- Retrieval: the user's question is converted into a vector, and the most semantically relevant text chunks are found in a vector database or document store;
- Augmentation: the retrieved chunks are combined with the question into the prompt — like handing the model reference material for an open-book exam;
- Generation: the model composes its answer from that material; good implementations also cite the source of each claim so it can be verified.
Note the boundary: RAG swaps the "reference material," not the model weights; the quality ceiling is set by the retrieved chunks.
RAG vs. Fine-Tuning
The one-line distinction: RAG changes the model's reference material; fine-tuning changes the model itself.
- Need the model to master company-internal documents, private data, or time-sensitive information → choose RAG: data stays fresh, costs stay low, answers are traceable;
- Need the model to change its tone, output format, or reason more like a domain expert → choose fine-tuning: the capability is written into the weights.
A common mistake is treating RAG as "fine-tuning for the poor": RAG trains nothing and can't fix behavior-level problems; conversely, fine-tuning can't fix knowledge freshness — the day after training finishes, new knowledge is already stale. To see how low-cost fine-tuning works, see What Is LoRA Fine-Tuning? Why Small Budgets Can Still Train Specialized Models.
RAG vs. Prompt Engineering
Prompt engineering answers "how to ask"; RAG answers "with what material to ask."
A well-crafted prompt helps the model better mobilize what it already knows and follow a required format — but it cannot conjure facts the model never learned. RAG delivers the facts directly. In real projects the two often work together: RAG supplies the material, the prompt sets the rules — "answer only from the provided material, say plainly when it isn't there, and cite the source of each piece of information."
Three Common Misconceptions
Misconception 1: RAG just means giving the model web search. Web search is only one special case of RAG (retrieval source: the whole internet). In enterprises, RAG more often retrieves from internal knowledge bases, ticket systems, or product documentation; the core of RAG is the "retrieval + generation" architecture, not the act of "going online."
Misconception 2: RAG eliminates hallucinations. It reduces them, it doesn't eradicate them: irrelevant retrieved chunks or a model misreading the material still produce errors. RAG's real value is making errors checkable — the answer carries its sources, so a mistake can be traced to the offending chunk.
Misconception 3: which is superior, RAG or fine-tuning? They are different dimensions: one governs "where knowledge comes from," the other "what the capability looks like." Production systems often use both: fine-tuning sets the behavioral style, RAG plugs in the latest knowledge.
Where It Fits Best
Enterprise knowledge-base Q&A (customer support, automatic answers in the internal wiki), long-document assistants (asking questions over hundred-page contracts, manuals, or papers), and any scenario where "a wrong answer is costly and sources must be verifiable." If you need to handle global questions spanning multiple documents (e.g., "what common trend do these financial reports reflect?"), take a look at its variant GraphRAG.