RAG vs Fine-Tuning vs Long Context
A decision framework for giving a model knowledge and behaviour, comparing RAG, fine-tuning and long context on freshness, cost, citations and control, plus how they combine.
Lesson 7 of 7 8 min
Three different tools
| RAG | Fine-tuning | Long context | |
|---|---|---|---|
| What it changes | What the model sees per request | The model's weights | How much the model sees per request |
| Best for | Facts, documents, fresh data | Behaviour, format, style, narrow tasks | Whole-document reasoning |
| Freshness | Update the index: minutes | Retrain: days | Per request: instant |
| Citations | Natural: you know the sources | None. Knowledge is diffused into weights | Possible, but less precise |
| Permissions | Filter per user | Anyone using the model gets everything trained in | Per request |
| Cost profile | Index infrastructure, plus a few K tokens per request | Training runs, eval, and hosting a custom model | Paying for many tokens every request |
| Hallucination on facts | Lower, since it's grounded | Still hallucinates. Can mix learned "facts" | Lower, but degrades with length |
Why fine-tuning is a poor way to add knowledge
It is tempting: "train the model on our docs so it knows them". In practice:
- The model learns style and patterns from the docs far more reliably than specific facts, and it still confidently invents details.
- You can't update a single fact without another training run, and you can't delete one at all, which matters for GDPR.
- You can't cite where an answer came from.
- Permissions don't exist. Everyone who can call the model can extract anything it learned.
Fine-tuning shines at behaviour:
- Always output this JSON schema or house style.
- Classify tickets into our 40 categories accurately.
- Use our domain vocabulary and abbreviations correctly.
- Distillation: make an 8B model match a frontier model's quality on one narrow task, at a small fraction of the cost. See Module 9.
Long context vs RAG
| Choose long context when | Choose RAG when |
|---|---|
| One big document per task (a contract, a codebase snapshot) | Large corpus (thousands to billions of chunks) |
| Low request volume | High request volume |
| Cross-references across the whole input matter | Answers live in a few passages |
| The same prefix repeats (cache it) | Content changes often and must be permission-filtered |
These are not exclusive. A common pattern is RAG to select documents, then long context to read them fully: retrieve the top 3 relevant documents and include them whole instead of in fragments.
Decision flow
Is it knowledge or behaviour?
Knowledge → RAG
+ long context for deep reads
Behaviour → prompt first
Still failing evals?
then fine-tune
Cost-bound at scale?
distil to a small model
- Try prompting with good instructions and a few examples. It is cheap and fast to iterate on.
- Add RAG if the model needs information it doesn't have.
- Fine-tune if the eval shows a behaviour gap prompting can't close, or if you need to cut cost or latency with a smaller model.
- Combine as needed. Fine-tuned models often do RAG, trained specifically to use retrieved context well and to cite.
Cost comparison sketch
For a support assistant at 1M requests a month:
- RAG: ≈ 5K input tokens per request (mostly retrieved context) on a mid-tier model, plus vector infrastructure.
- Long context (whole 300K-token manual per request): 60× more input tokens per request. Even with prompt caching at about 10% of the price, that is roughly 6× the input cost, with higher TTFT.
- Fine-tuned small model with RAG: fewer tokens (shorter instructions, since behaviour is trained in) and a model several times cheaper per token, plus training and hosting costs.
Key takeaways
- Use RAG for knowledge that changes, must be cited, or is permission-scoped.
- Use fine-tuning for behaviour, such as format, style, domain language or narrow tasks, and for distilling a big model into a cheaper one.
- Use long context for occasional deep analysis of a single large input, or a stable prefix that can be prompt-cached.
- They combine. A fine-tuned model can answer over retrieved context, inside a long window.
Go deeper
Finished reading? Mark it done to track your progress.