The Unit Economics of Tokens
How to think about LLM cost per request, per user and per outcome, what drives token prices, and why unit economics decide which AI features are viable.
A new kind of cost structure
Classic SaaS has near-zero marginal cost per request. Once servers are paid for, one more page view costs almost nothing. LLM features break that: every use costs real money, often cents, and heavy users can cost more than they pay.
That makes unit economics a design input, not an afterthought.
The formula, extended
Cost per request = (fresh input × input price) + (cached input × cached price) + (output × output price) + retrieval + tools + guardrails
Example: a writing assistant on a mid-tier model ($2 / $10 per million):
- Per request: 3K input + 600 output tokens = $0.006 + $0.006 = $0.012
- Typical user: 15 requests a day × 22 days = 330 requests = ≈ $4 a month
- Power user (95th percentile): 150 requests a day = ≈ $40 a month
On a $20-a-month plan, the typical user leaves a healthy margin, and the power user costs twice their subscription. That shapes pricing (usage tiers, fair-use limits) and architecture (routing heavy users' simple requests to cheaper models, caching aggressively).
What drives the price per token
| Driver | Typical spread |
|---|---|
| Model tier: small vs frontier | 20–100× |
| Output vs input | Output ≈ 4–8× input |
| Cached vs fresh input | Cached ≈ 10% of fresh |
| Batch vs real-time | Batch ≈ 50% off |
| Reasoning / thinking tokens | Billed as output; can multiply output 2–10× |
| Long-context surcharges | Some models charge more above a context threshold |
| Self-hosted vs API | Depends on utilisation (build vs buy) |
Token prices for a given capability level have been falling quickly, often severalfold per year. Design for today's prices, but know that a feature too expensive now may be viable in 12 months, and plan model upgrades accordingly.
Cost per outcome
The most useful metric is cost per successful outcome:
- Cost per resolved support ticket, not per message
- Cost per accepted code suggestion
- Cost per qualified lead enriched
A frontier model at 3× the price that resolves 90% of tickets instead of 60% may be cheaper per resolution, and much cheaper once you count human handling time for the failures.
Distribution, not average
Usage is heavy-tailed. The top 5–10% of users often drive 50% or more of token spend. Plan for it:
- Per-user limits and fair-use policies by plan
- Tiered models: default to efficient models, and use premium models on request or for paid tiers
- Monitoring cost per user percentiles, not only the total
- Abuse detection: scripted use of free tiers can burn thousands of dollars a day
Key takeaways
- LLM cost is variable per use. Cost per request = input tokens × input price + output tokens × output price, plus retrieval and tool costs.
- Express cost per user per month and per successful outcome, and compare it with revenue or value per user.
- Price drivers include model tier, input vs output, cached vs fresh input, batch vs real-time, and reasoning tokens.
- Heavy users dominate cost. Design plans, limits and routing around the usage distribution, not the average.