API Costs at Scale
How hosted-API costs behave at high volume, covering rate-limit tiers, provisioned throughput, caching and batch discounts, committed-use deals, and a playbook for cutting the bill.
Lesson 2 of 5 9 min
What changes at scale
At small scale, API cost is simply tokens × list price. At millions of requests a day, new concerns appear:
- Rate limits: default tiers cap tokens and requests per minute, and your peak may exceed them. Upgrades usually depend on spend history, so request them well before launch.
- Latency variance: shared, pay-as-you-go capacity has noisier latency at busy times.
- Capacity guarantees: during demand spikes, providers may throttle pay-as-you-go traffic first.
- Negotiation: at large spend, enterprise agreements, committed-use discounts and custom terms become possible.
Pricing options
| Option | How you pay | Best for |
|---|---|---|
| Pay-as-you-go | Per token, list price | Variable or early-stage traffic |
| Prompt caching | Discounted cached input (≈ 10%), sometimes a cache-write premium | Repeated prefixes: system prompts, documents, chat history |
| Batch API | ≈ 50% off, asynchronous | Non-interactive bulk work |
| Provisioned throughput | Fixed hourly price for reserved capacity | Steady high volume, strict latency SLOs |
| Committed-use / enterprise | Discounts for spend commitments | Large, predictable spend |
Provisioned throughput
Instead of paying per token, you reserve model capacity, measured in throughput units, for a fixed hourly or monthly fee.
- Pros: guaranteed capacity, more consistent latency, and a lower effective price per token at high utilisation.
- Cons: you pay whether you use it or not, which is the same utilisation problem as self-hosting, and it often requires a commitment.
- Pattern: provision for your baseline load and send spikes to pay-as-you-go ("burst"). Size the reservation from your p50 traffic, not your peak.
Worked example: cutting a bill
Starting point: 10M requests a day, 2K input and 400 output tokens, on a $2 / $10 model, which is $2.4M a month at list price (back-of-envelope lesson).
| Step | Change | New monthly cost |
|---|---|---|
| 0 | Baseline | $2.40M |
| 1 | Prompt caching: 1.2K of the 2K input tokens are a stable prefix, cached at 10% | Input $1.2M → $0.55M, total $1.75M |
| 2 | Route 60% of traffic to a small model ($0.25 / $2) | ≈ $0.9M |
| 3 | Move 20% of volume (summaries, tagging) to the batch API | ≈ $0.8M |
| 4 | Trim output: concise format, 400 → 300 tokens | ≈ $0.65M |
That is roughly a 3.5× reduction with no self-hosting and the same frontier-quality path for the hard 40%. Each step needs an eval to confirm quality holds.
Monitoring spend
- Dashboards by feature, tenant and model. Unexplained cost is usually one feature or one customer.
- Budget alerts at daily granularity. Runaway agents and abusive traffic can burn a month's budget in a day.
- Cost per outcome trends. Rising cost with flat outcomes signals prompt bloat or context growth.
- Cache hit rate: a sudden drop is an immediate cost spike.
Key takeaways
- At scale, the constraint is often rate limits and capacity guarantees, not just price. Plan quota tiers and provisioned throughput early.
- Caching and batch discounts routinely cut API bills by 30–70% without changing models.
- Provisioned throughput (reserved capacity) trades a fixed hourly fee for predictable latency and guaranteed capacity. It pays off at steady, high utilisation.
- Cut costs in order. Use fewer tokens, then cheaper tokens, then negotiated pricing, and self-host only after that.
Go deeper
Finished reading? Mark it done to track your progress.