Multi-Agent Architectures
When splitting work across several agents helps, the main patterns (orchestrator-worker, pipelines, handoffs, debate), and their costs in tokens, latency and debuggability.
Why more than one agent?
A single agent with 40 tools and a 150K-token history gets slow, expensive and confused. Splitting the work helps because of:
- Context isolation: a sub-agent researching one question works in a fresh context. It returns a 500-token summary instead of adding 30K tokens of search results to the main agent's history.
- Specialisation: each agent gets a focused prompt and a small tool set, such as "SQL analyst" or "code reviewer".
- Parallelism: independent subtasks run at the same time, which cuts wall-clock time.
- Different models: a frontier model plans, and cheaper models do the bulk work.
Patterns
Orchestrator-worker
A lead agent breaks the task down, spawns workers (often in parallel) with specific instructions, and synthesises their results. This suits research, broad code changes, and due-diligence tasks. It is the most widely useful multi-agent pattern.
Pipeline (assembly line)
Agents in a fixed sequence, each transforming the output of the previous one: planner → implementer → tester → reviewer. It is essentially a workflow whose steps happen to be agents. It is predictable and easy to evaluate stage by stage.
Handoff (routing)
A front agent triages and transfers the conversation to a specialist, such as billing, technical support or sales, each with its own tools and permissions. This is common in customer-support systems. The key design question is what context passes along at handoff.
Critic / evaluator loop
One agent produces and another critiques, repeating until the critic approves or a limit is reached. It improves quality on code, writing and plans, at 2–3× the tokens.
Debate / ensemble
Several agents answer independently, then vote or argue toward a consensus. It is expensive, so reserve it for high-stakes decisions where diversity of reasoning reduces errors.
Costs and risks
- Coordination failures: workers duplicate effort, misread instructions, or return results in inconsistent formats. Give workers precise task specs and structured output schemas.
- Error propagation: a worker's confident mistake gets synthesised into the final answer. The orchestrator should check claims, and ask for evidence and citations.
- Debugging: you need end-to-end tracing across agents (parent and child spans, prompts and tool calls) to see what happened. See tracing.
- Shared state: agents writing to the same resources (files, records) need locking or clear ownership.
When to use a single agent instead
- The task is sequential and tightly coupled. Each step depends on the last, so there is nothing to parallelise.
- The context fits comfortably in one window.
- Latency and cost matter more than a small quality gain.
Start with one agent and good tools. Add sub-agents when you hit a specific limit, such as context overflow, obvious parallelism, or a specialist tool set, and confirm the gain with evals.
Key takeaways
- Multi-agent systems split a task across LLM instances with separate contexts, prompts and tools.
- The main win is context isolation. Each sub-agent works in a clean, focused context and returns a compact result.
- Orchestrator-worker is the most useful pattern. Pipelines, handoffs and critic loops cover most other cases.
- Costs multiply. Several agents means several times the tokens, and failures become harder to trace. Use them only where the eval shows a gain.