GenAI System Design
4. RAG Architecture

GraphRAG and Agentic Retrieval

When single-shot vector retrieval isn't enough, and what to use instead: knowledge graphs, GraphRAG summaries, structured and SQL retrieval, and agents that search iteratively.

Lesson 6 of 7 10 min

Where single-shot RAG breaks

Standard RAG retrieves once and answers once. That works for lookup questions: "What is the refund window for annual plans?" It fails for:

Question typeExampleWhy vector RAG struggles
Multi-hop"Which team owns the service that caused last week's outage?"Needs outage report → service → owner, found in different documents
Aggregate"How many P1 incidents did we have per quarter?"Needs counting over many records, not similar passages
Global / thematic"What are the main complaints across all customer feedback?"No single chunk contains the answer. It needs the whole corpus
Structured"Top 10 customers by revenue in EMEA"The data lives in a database, not in text

Knowledge graphs and GraphRAG

A knowledge graph stores entities (people, services, products) as nodes and relationships (owns, depends-on, reported) as edges. Multi-hop questions become graph traversals.

GraphRAG, popularised by Microsoft Research, builds this automatically with LLMs:

Chunks
LLM extracts
entities + relations
Graph
merge duplicates
Communities
cluster related entities
LLM summaries
per community, at multiple levels

At query time:

  • Local search: find entities in the question, walk their neighbourhood, and gather connected facts and source chunks. Good for multi-hop.
  • Global search: answer from the community summaries, map-reduce style. Good for "main themes across everything".

Trade-offs:

  • Ingestion cost: an LLM reads every chunk to extract entities, then writes summaries. This can cost 10–100× more than embedding the same corpus, and must be partly redone as content changes.
  • Extraction quality: entity resolution (is "AWS" the same as "Amazon Web Services"?) and missed relations directly limit answers.
  • Best fit: relatively stable corpora with many connected entities, such as investigations, research, legal discovery or organisational knowledge.

A lighter-weight alternative: if you already have structured relationships (a service catalogue, CRM, org chart), query them directly instead of re-extracting them from text with an LLM.

Structured retrieval

When the answer lives in tables, retrieve it with queries:

  • Text-to-SQL: the LLM writes SQL against a described schema, and you execute it read-only with row limits. The full design problem is covered in the text-to-SQL design problem.
  • API tools: call internal APIs (orders, billing, inventory) with the user's credentials.
  • Metadata filters extracted from the question, combined with vector search.

Never embed rows of a large table and hope vector search does arithmetic.

Agentic retrieval

Give the model search as a tool and let it decide what to look up:

  1. Search for "last week's outage report".
  2. Read it and find the service name.
  3. Search the service catalogue for the owner.
  4. Answer with citations to both.
Single-shot RAGAgentic retrieval
Latency≈ 1–3 s to first tokenSeveral seconds to minutes
Cost1 LLM call3–20+ calls, growing context
HandlesLookupMulti-hop, exploratory, cross-source
PredictabilityHighLower. Needs step limits and evals

A practical design is a router: classify the question, send simple lookups down the fast RAG path, and send complex ones to an agent with a step and token budget. Module 5 covers agent loops.

Key takeaways

  • Plain RAG answers "find the passage" questions well. It struggles with multi-hop, aggregate and global questions.
  • GraphRAG extracts entities and relations into a graph and pre-computes community summaries, which helps with connected and corpus-wide questions, at a large ingestion cost.
  • Structured data belongs in structured retrieval. Use text-to-SQL or APIs, not embeddings of table rows.
  • Agentic retrieval lets the model search, read and search again. It is more capable, but slower, costlier and harder to bound.

Go deeper