GraphRAG and Agentic Retrieval
When single-shot vector retrieval isn't enough, and what to use instead: knowledge graphs, GraphRAG summaries, structured and SQL retrieval, and agents that search iteratively.
Where single-shot RAG breaks
Standard RAG retrieves once and answers once. That works for lookup questions: "What is the refund window for annual plans?" It fails for:
| Question type | Example | Why vector RAG struggles |
|---|---|---|
| Multi-hop | "Which team owns the service that caused last week's outage?" | Needs outage report → service → owner, found in different documents |
| Aggregate | "How many P1 incidents did we have per quarter?" | Needs counting over many records, not similar passages |
| Global / thematic | "What are the main complaints across all customer feedback?" | No single chunk contains the answer. It needs the whole corpus |
| Structured | "Top 10 customers by revenue in EMEA" | The data lives in a database, not in text |
Knowledge graphs and GraphRAG
A knowledge graph stores entities (people, services, products) as nodes and relationships (owns, depends-on, reported) as edges. Multi-hop questions become graph traversals.
GraphRAG, popularised by Microsoft Research, builds this automatically with LLMs:
At query time:
- Local search: find entities in the question, walk their neighbourhood, and gather connected facts and source chunks. Good for multi-hop.
- Global search: answer from the community summaries, map-reduce style. Good for "main themes across everything".
Trade-offs:
- Ingestion cost: an LLM reads every chunk to extract entities, then writes summaries. This can cost 10–100× more than embedding the same corpus, and must be partly redone as content changes.
- Extraction quality: entity resolution (is "AWS" the same as "Amazon Web Services"?) and missed relations directly limit answers.
- Best fit: relatively stable corpora with many connected entities, such as investigations, research, legal discovery or organisational knowledge.
A lighter-weight alternative: if you already have structured relationships (a service catalogue, CRM, org chart), query them directly instead of re-extracting them from text with an LLM.
Structured retrieval
When the answer lives in tables, retrieve it with queries:
- Text-to-SQL: the LLM writes SQL against a described schema, and you execute it read-only with row limits. The full design problem is covered in the text-to-SQL design problem.
- API tools: call internal APIs (orders, billing, inventory) with the user's credentials.
- Metadata filters extracted from the question, combined with vector search.
Never embed rows of a large table and hope vector search does arithmetic.
Agentic retrieval
Give the model search as a tool and let it decide what to look up:
- Search for "last week's outage report".
- Read it and find the service name.
- Search the service catalogue for the owner.
- Answer with citations to both.
| Single-shot RAG | Agentic retrieval | |
|---|---|---|
| Latency | ≈ 1–3 s to first token | Several seconds to minutes |
| Cost | 1 LLM call | 3–20+ calls, growing context |
| Handles | Lookup | Multi-hop, exploratory, cross-source |
| Predictability | High | Lower. Needs step limits and evals |
A practical design is a router: classify the question, send simple lookups down the fast RAG path, and send complex ones to an agent with a step and token budget. Module 5 covers agent loops.
Key takeaways
- Plain RAG answers "find the passage" questions well. It struggles with multi-hop, aggregate and global questions.
- GraphRAG extracts entities and relations into a graph and pre-computes community summaries, which helps with connected and corpus-wide questions, at a large ingestion cost.
- Structured data belongs in structured retrieval. Use text-to-SQL or APIs, not embeddings of table rows.
- Agentic retrieval lets the model search, read and search again. It is more capable, but slower, costlier and harder to bound.