GenAI System Design
8. Cost and Capacity Planning

The Unit Economics of Tokens

How to think about LLM cost per request, per user and per outcome, what drives token prices, and why unit economics decide which AI features are viable.

Lesson 1 of 5 8 min

A new kind of cost structure

Classic SaaS has near-zero marginal cost per request. Once servers are paid for, one more page view costs almost nothing. LLM features break that: every use costs real money, often cents, and heavy users can cost more than they pay.

That makes unit economics a design input, not an afterthought.

The formula, extended

Cost per request = (fresh input × input price) + (cached input × cached price) + (output × output price) + retrieval + tools + guardrails

Cost per request
× requests per session
× sessions per user per month
= Cost per user / month
vs revenue per user
gross margin

Example: a writing assistant on a mid-tier model ($2 / $10 per million):

  • Per request: 3K input + 600 output tokens = $0.006 + $0.006 = $0.012
  • Typical user: 15 requests a day × 22 days = 330 requests = ≈ $4 a month
  • Power user (95th percentile): 150 requests a day = ≈ $40 a month

On a $20-a-month plan, the typical user leaves a healthy margin, and the power user costs twice their subscription. That shapes pricing (usage tiers, fair-use limits) and architecture (routing heavy users' simple requests to cheaper models, caching aggressively).

What drives the price per token

DriverTypical spread
Model tier: small vs frontier20–100×
Output vs inputOutput ≈ 4–8× input
Cached vs fresh inputCached ≈ 10% of fresh
Batch vs real-timeBatch ≈ 50% off
Reasoning / thinking tokensBilled as output; can multiply output 2–10×
Long-context surchargesSome models charge more above a context threshold
Self-hosted vs APIDepends on utilisation (build vs buy)

Token prices for a given capability level have been falling quickly, often severalfold per year. Design for today's prices, but know that a feature too expensive now may be viable in 12 months, and plan model upgrades accordingly.

Cost per outcome

The most useful metric is cost per successful outcome:

  • Cost per resolved support ticket, not per message
  • Cost per accepted code suggestion
  • Cost per qualified lead enriched

A frontier model at 3× the price that resolves 90% of tickets instead of 60% may be cheaper per resolution, and much cheaper once you count human handling time for the failures.

Distribution, not average

Usage is heavy-tailed. The top 5–10% of users often drive 50% or more of token spend. Plan for it:

  • Per-user limits and fair-use policies by plan
  • Tiered models: default to efficient models, and use premium models on request or for paid tiers
  • Monitoring cost per user percentiles, not only the total
  • Abuse detection: scripted use of free tiers can burn thousands of dollars a day

Key takeaways

  • LLM cost is variable per use. Cost per request = input tokens × input price + output tokens × output price, plus retrieval and tool costs.
  • Express cost per user per month and per successful outcome, and compare it with revenue or value per user.
  • Price drivers include model tier, input vs output, cached vs fresh input, batch vs real-time, and reasoning tokens.
  • Heavy users dominate cost. Design plans, limits and routing around the usage distribution, not the average.

Go deeper

Finished reading? Mark it done to track your progress.