GenAI System Design
5. Agents and Tool Use

Controlling Agent Cost and Latency

Why agent costs grow faster than linearly with steps, how to model them, and the levers that keep agents affordable and fast, including caching, trimming, routing, budgets and parallelism.

Lesson 7 of 7 9 min

Why agents are expensive

A chat turn is one LLM call. An agent task is 5–50 calls, and each call resends everything before it: system prompt, tool definitions, every previous tool call and result. If context grows by g tokens per step from a base B, step n sends B + (n − 1)g tokens. The total over N steps is about N·B + g·N²/2, which is quadratic in steps.

How agent context and cost grow

Every step resends the whole conversation so far: system prompt, tool definitions, and every previous tool call and result.

Earlier context billed at the cached rate

Keep ≤ 300 tokens per result

Cumulative cost of one task
$0.00$0.11$0.2211120step
Cost per task
$0.22
Input tokens processed
462K
3.9× a non-growing context
Final context size
40.2K tokens
4% of the window
At 10K tasks a day
$67,428/mo

Total input grows roughly with the square of the step count, because step n resends everything from steps 1 to n−1. Caching and trimming tool output are the two biggest levers. Capping steps is the third.

Modelling cost per task

Estimate these from traces or a pilot:

  • Steps per task, as a distribution. Use p50 and p90, not the mean.
  • Base prompt: system prompt plus tool definitions.
  • Growth per step: tool result tokens plus model output tokens.
  • Model prices, including cached input and cache writes.

Then cost per task × tasks per day gives the budget. Present it alongside the business value of the task. $0.50 to resolve a support ticket is cheap, and $0.50 to answer an FAQ is not.

The levers

LeverEffectNotes
Prompt cachingCached prefix billed at ≈ 10%Keep the prefix stable: tools and system prompt first, no timestamps up top
Trim tool resultsSmaller growth gReturn the relevant fields or lines. Let the agent fetch more if needed
CompactionResets context sizeSummarise history every N steps, and keep key facts and the plan
Sub-agentsIsolate big explorationsThe main context only gets the compact result
Model routingCheaper tokensSmall model for simple steps (formatting, extraction), frontier for planning
Fewer stepsLess of everythingBetter tools (one search_and_summarise instead of search + fetch + read), better prompts, examples
Parallel tool callsLower latencyIndependent lookups in one turn
BudgetsBounded worst caseHard caps on steps, tokens, $ and time per task

Latency

Each step costs TTFT + generation time + tool execution time. With 15 steps at about 3 s each, a task takes 45 s before any tool slowness.

  • Parallelise independent calls and sub-agents.
  • Stream progress to the user ("Searching orders…", "Found 3 matches…") so long tasks feel alive.
  • Run asynchronously for long tasks: return a task ID and notify on completion instead of holding a request open.
  • Speculate cheaply: start likely-needed lookups before the model asks for them, for example the user's account details.
  • Use smaller, faster models for steps that don't need deep reasoning.

Budgets as product design

Budgets are a product decision as well as a safety net:

  • Per task: "at most 20 steps or $1". When the limit hits, return the partial result with an explanation, not an error.
  • Per user or tenant per day: protects against abuse and runaway loops.
  • Alerting: on cost per task p95 and on tasks that hit their limits, which are often bugs or stuck loops.

Key takeaways

  • Each step resends the whole growing context, so total input tokens grow roughly with the square of the step count.
  • Prompt caching is the biggest single lever, because most of each step's input is an unchanged prefix.
  • Trim tool outputs, compact history, and delegate to sub-agents to keep the context small.
  • Enforce per-task budgets for steps, tokens, dollars and time. Use cheaper models for easy steps and run independent calls in parallel.

Go deeper

Finished reading? Mark it done to track your progress.