GenAI System Design
8. Cost and Capacity Planning

API Costs at Scale

How hosted-API costs behave at high volume, covering rate-limit tiers, provisioned throughput, caching and batch discounts, committed-use deals, and a playbook for cutting the bill.

Lesson 2 of 5 9 min

What changes at scale

At small scale, API cost is simply tokens × list price. At millions of requests a day, new concerns appear:

  • Rate limits: default tiers cap tokens and requests per minute, and your peak may exceed them. Upgrades usually depend on spend history, so request them well before launch.
  • Latency variance: shared, pay-as-you-go capacity has noisier latency at busy times.
  • Capacity guarantees: during demand spikes, providers may throttle pay-as-you-go traffic first.
  • Negotiation: at large spend, enterprise agreements, committed-use discounts and custom terms become possible.

Pricing options

OptionHow you payBest for
Pay-as-you-goPer token, list priceVariable or early-stage traffic
Prompt cachingDiscounted cached input (≈ 10%), sometimes a cache-write premiumRepeated prefixes: system prompts, documents, chat history
Batch API≈ 50% off, asynchronousNon-interactive bulk work
Provisioned throughputFixed hourly price for reserved capacitySteady high volume, strict latency SLOs
Committed-use / enterpriseDiscounts for spend commitmentsLarge, predictable spend

Provisioned throughput

Instead of paying per token, you reserve model capacity, measured in throughput units, for a fixed hourly or monthly fee.

  • Pros: guaranteed capacity, more consistent latency, and a lower effective price per token at high utilisation.
  • Cons: you pay whether you use it or not, which is the same utilisation problem as self-hosting, and it often requires a commitment.
  • Pattern: provision for your baseline load and send spikes to pay-as-you-go ("burst"). Size the reservation from your p50 traffic, not your peak.

Worked example: cutting a bill

Starting point: 10M requests a day, 2K input and 400 output tokens, on a $2 / $10 model, which is $2.4M a month at list price (back-of-envelope lesson).

StepChangeNew monthly cost
0Baseline$2.40M
1Prompt caching: 1.2K of the 2K input tokens are a stable prefix, cached at 10%Input $1.2M → $0.55M, total $1.75M
2Route 60% of traffic to a small model ($0.25 / $2)≈ $0.9M
3Move 20% of volume (summaries, tagging) to the batch API≈ $0.8M
4Trim output: concise format, 400 → 300 tokens≈ $0.65M

That is roughly a 3.5× reduction with no self-hosting and the same frontier-quality path for the hard 40%. Each step needs an eval to confirm quality holds.

Monitoring spend

  • Dashboards by feature, tenant and model. Unexplained cost is usually one feature or one customer.
  • Budget alerts at daily granularity. Runaway agents and abusive traffic can burn a month's budget in a day.
  • Cost per outcome trends. Rising cost with flat outcomes signals prompt bloat or context growth.
  • Cache hit rate: a sudden drop is an immediate cost spike.

Key takeaways

  • At scale, the constraint is often rate limits and capacity guarantees, not just price. Plan quota tiers and provisioned throughput early.
  • Caching and batch discounts routinely cut API bills by 30–70% without changing models.
  • Provisioned throughput (reserved capacity) trades a fixed hourly fee for predictable latency and guaranteed capacity. It pays off at steady, high utilisation.
  • Cut costs in order. Use fewer tokens, then cheaper tokens, then negotiated pricing, and self-host only after that.

Go deeper

Finished reading? Mark it done to track your progress.