Controlling Agent Cost and Latency
Why agent costs grow faster than linearly with steps, how to model them, and the levers that keep agents affordable and fast, including caching, trimming, routing, budgets and parallelism.
Why agents are expensive
A chat turn is one LLM call. An agent task is 5–50 calls, and each call resends everything before it: system prompt, tool definitions, every previous tool call and result. If context grows by g tokens per step from a base B, step n sends B + (n − 1)g tokens. The total over N steps is about N·B + g·N²/2, which is quadratic in steps.
How agent context and cost grow
Every step resends the whole conversation so far: system prompt, tool definitions, and every previous tool call and result.
Earlier context billed at the cached rate
Keep ≤ 300 tokens per result
Total input grows roughly with the square of the step count, because step n resends everything from steps 1 to n−1. Caching and trimming tool output are the two biggest levers. Capping steps is the third.
Modelling cost per task
Estimate these from traces or a pilot:
- Steps per task, as a distribution. Use p50 and p90, not the mean.
- Base prompt: system prompt plus tool definitions.
- Growth per step: tool result tokens plus model output tokens.
- Model prices, including cached input and cache writes.
Then cost per task × tasks per day gives the budget. Present it alongside the business value of the task. $0.50 to resolve a support ticket is cheap, and $0.50 to answer an FAQ is not.
The levers
| Lever | Effect | Notes |
|---|---|---|
| Prompt caching | Cached prefix billed at ≈ 10% | Keep the prefix stable: tools and system prompt first, no timestamps up top |
| Trim tool results | Smaller growth g | Return the relevant fields or lines. Let the agent fetch more if needed |
| Compaction | Resets context size | Summarise history every N steps, and keep key facts and the plan |
| Sub-agents | Isolate big explorations | The main context only gets the compact result |
| Model routing | Cheaper tokens | Small model for simple steps (formatting, extraction), frontier for planning |
| Fewer steps | Less of everything | Better tools (one search_and_summarise instead of search + fetch + read), better prompts, examples |
| Parallel tool calls | Lower latency | Independent lookups in one turn |
| Budgets | Bounded worst case | Hard caps on steps, tokens, $ and time per task |
Latency
Each step costs TTFT + generation time + tool execution time. With 15 steps at about 3 s each, a task takes 45 s before any tool slowness.
- Parallelise independent calls and sub-agents.
- Stream progress to the user ("Searching orders…", "Found 3 matches…") so long tasks feel alive.
- Run asynchronously for long tasks: return a task ID and notify on completion instead of holding a request open.
- Speculate cheaply: start likely-needed lookups before the model asks for them, for example the user's account details.
- Use smaller, faster models for steps that don't need deep reasoning.
Budgets as product design
Budgets are a product decision as well as a safety net:
- Per task: "at most 20 steps or $1". When the limit hits, return the partial result with an explanation, not an error.
- Per user or tenant per day: protects against abuse and runaway loops.
- Alerting: on cost per task p95 and on tasks that hit their limits, which are often bugs or stuck loops.
Key takeaways
- Each step resends the whole growing context, so total input tokens grow roughly with the square of the step count.
- Prompt caching is the biggest single lever, because most of each step's input is an unchanged prefix.
- Trim tool outputs, compact history, and delegate to sub-agents to keep the context small.
- Enforce per-task budgets for steps, tokens, dollars and time. Use cheaper models for easy steps and run independent calls in parallel.