GenAI System Design
4. RAG Architecture

RAG vs Fine-Tuning vs Long Context

A decision framework for giving a model knowledge and behaviour, comparing RAG, fine-tuning and long context on freshness, cost, citations and control, plus how they combine.

Lesson 7 of 7 8 min

Three different tools

RAGFine-tuningLong context
What it changesWhat the model sees per requestThe model's weightsHow much the model sees per request
Best forFacts, documents, fresh dataBehaviour, format, style, narrow tasksWhole-document reasoning
FreshnessUpdate the index: minutesRetrain: daysPer request: instant
CitationsNatural: you know the sourcesNone. Knowledge is diffused into weightsPossible, but less precise
PermissionsFilter per userAnyone using the model gets everything trained inPer request
Cost profileIndex infrastructure, plus a few K tokens per requestTraining runs, eval, and hosting a custom modelPaying for many tokens every request
Hallucination on factsLower, since it's groundedStill hallucinates. Can mix learned "facts"Lower, but degrades with length

Why fine-tuning is a poor way to add knowledge

It is tempting: "train the model on our docs so it knows them". In practice:

  • The model learns style and patterns from the docs far more reliably than specific facts, and it still confidently invents details.
  • You can't update a single fact without another training run, and you can't delete one at all, which matters for GDPR.
  • You can't cite where an answer came from.
  • Permissions don't exist. Everyone who can call the model can extract anything it learned.

Fine-tuning shines at behaviour:

  • Always output this JSON schema or house style.
  • Classify tickets into our 40 categories accurately.
  • Use our domain vocabulary and abbreviations correctly.
  • Distillation: make an 8B model match a frontier model's quality on one narrow task, at a small fraction of the cost. See Module 9.

Long context vs RAG

Choose long context whenChoose RAG when
One big document per task (a contract, a codebase snapshot)Large corpus (thousands to billions of chunks)
Low request volumeHigh request volume
Cross-references across the whole input matterAnswers live in a few passages
The same prefix repeats (cache it)Content changes often and must be permission-filtered

These are not exclusive. A common pattern is RAG to select documents, then long context to read them fully: retrieve the top 3 relevant documents and include them whole instead of in fragments.

Decision flow

Is it knowledge or behaviour?
Knowledge → RAG
+ long context for deep reads
Behaviour → prompt first
Still failing evals?
then fine-tune
Cost-bound at scale?
distil to a small model
  1. Try prompting with good instructions and a few examples. It is cheap and fast to iterate on.
  2. Add RAG if the model needs information it doesn't have.
  3. Fine-tune if the eval shows a behaviour gap prompting can't close, or if you need to cut cost or latency with a smaller model.
  4. Combine as needed. Fine-tuned models often do RAG, trained specifically to use retrieved context well and to cite.

Cost comparison sketch

For a support assistant at 1M requests a month:

  • RAG: ≈ 5K input tokens per request (mostly retrieved context) on a mid-tier model, plus vector infrastructure.
  • Long context (whole 300K-token manual per request): 60× more input tokens per request. Even with prompt caching at about 10% of the price, that is roughly 6× the input cost, with higher TTFT.
  • Fine-tuned small model with RAG: fewer tokens (shorter instructions, since behaviour is trained in) and a model several times cheaper per token, plus training and hosting costs.

Key takeaways

  • Use RAG for knowledge that changes, must be cited, or is permission-scoped.
  • Use fine-tuning for behaviour, such as format, style, domain language or narrow tasks, and for distilling a big model into a cheaper one.
  • Use long context for occasional deep analysis of a single large input, or a stable prefix that can be prompt-cached.
  • They combine. A fine-tuned model can answer over retrieved context, inside a long window.

Go deeper

Finished reading? Mark it done to track your progress.