GenAI System Design
5. Agents and Tool Use

Multi-Agent Architectures

When splitting work across several agents helps, the main patterns (orchestrator-worker, pipelines, handoffs, debate), and their costs in tokens, latency and debuggability.

Lesson 4 of 7 9 min

Why more than one agent?

A single agent with 40 tools and a 150K-token history gets slow, expensive and confused. Splitting the work helps because of:

  • Context isolation: a sub-agent researching one question works in a fresh context. It returns a 500-token summary instead of adding 30K tokens of search results to the main agent's history.
  • Specialisation: each agent gets a focused prompt and a small tool set, such as "SQL analyst" or "code reviewer".
  • Parallelism: independent subtasks run at the same time, which cuts wall-clock time.
  • Different models: a frontier model plans, and cheaper models do the bulk work.

Patterns

Orchestrator-worker

Orchestrator
plans, delegates, synthesises
Worker A
search topic 1
Worker B
search topic 2
Worker C
analyse data
Synthesis
orchestrator combines results → answer
Workers run in parallel with their own contexts and return compact results.

A lead agent breaks the task down, spawns workers (often in parallel) with specific instructions, and synthesises their results. This suits research, broad code changes, and due-diligence tasks. It is the most widely useful multi-agent pattern.

Pipeline (assembly line)

Agents in a fixed sequence, each transforming the output of the previous one: planner → implementer → tester → reviewer. It is essentially a workflow whose steps happen to be agents. It is predictable and easy to evaluate stage by stage.

Handoff (routing)

A front agent triages and transfers the conversation to a specialist, such as billing, technical support or sales, each with its own tools and permissions. This is common in customer-support systems. The key design question is what context passes along at handoff.

Critic / evaluator loop

One agent produces and another critiques, repeating until the critic approves or a limit is reached. It improves quality on code, writing and plans, at 2–3× the tokens.

Debate / ensemble

Several agents answer independently, then vote or argue toward a consensus. It is expensive, so reserve it for high-stakes decisions where diversity of reasoning reduces errors.

Costs and risks

  • Coordination failures: workers duplicate effort, misread instructions, or return results in inconsistent formats. Give workers precise task specs and structured output schemas.
  • Error propagation: a worker's confident mistake gets synthesised into the final answer. The orchestrator should check claims, and ask for evidence and citations.
  • Debugging: you need end-to-end tracing across agents (parent and child spans, prompts and tool calls) to see what happened. See tracing.
  • Shared state: agents writing to the same resources (files, records) need locking or clear ownership.

When to use a single agent instead

  • The task is sequential and tightly coupled. Each step depends on the last, so there is nothing to parallelise.
  • The context fits comfortably in one window.
  • Latency and cost matter more than a small quality gain.

Start with one agent and good tools. Add sub-agents when you hit a specific limit, such as context overflow, obvious parallelism, or a specialist tool set, and confirm the gain with evals.

Key takeaways

  • Multi-agent systems split a task across LLM instances with separate contexts, prompts and tools.
  • The main win is context isolation. Each sub-agent works in a clean, focused context and returns a compact result.
  • Orchestrator-worker is the most useful pattern. Pipelines, handoffs and critic loops cover most other cases.
  • Costs multiply. Several agents means several times the tokens, and failures become harder to trace. Use them only where the eval shows a gain.

Go deeper