GenAI System Design
8. Cost and Capacity Planning

Self-Hosting Costs: GPU-Hours

What self-hosting really costs, including GPU pricing models, utilisation, cost per million tokens from throughput, the hidden costs, and how to raise utilisation.

Lesson 3 of 5 9 min

From GPU-hours to cost per token

Cost per 1M output tokens = (GPU $/hour × GPUs per replica) ÷ (replica tokens/s × 3,600 × utilisation) × 1,000,000

Example: Llama 3.1 70B in FP8 on 2 × H100 at $2.50 per GPU-hour, with ≈ 2,500 output tokens/s per replica (from the capacity planning lesson):

  • Replica cost: $5 an hour
  • At 100% utilisation: 2,500 × 3,600 = 9M tokens an hour, so ≈ $0.56 per million output tokens
  • At 40% average utilisation: ≈ $1.39 per million
  • At 15% utilisation (low traffic with redundancy): ≈ $3.70 per million

Compare that with API output prices for similar-quality models to see which side of the line you're on. Input (prefill) tokens are much cheaper to process per token than output, so for a like-for-like comparison, count both.

GPU pricing models

ModelTypical discount vs on-demandTrade-off
On-demand cloud—Flexible, but availability of top GPUs isn't guaranteed
Reserved / committed (1–3 years)Large, often 30–60%Locked in as prices fall and hardware improves
Spot / preemptibleVery largeCan be reclaimed at short notice. Suits batch, not serving
Specialised GPU cloudsOften cheaper than hyperscalersFewer surrounding services, varying reliability
Owned hardwareLowest at high utilisation over yearsCapex, data-centre power and cooling, operations, refresh cycles

Utilisation: the real driver

You provision for peak plus redundancy, but pay for every hour:

  • Diurnal patterns: traffic may be 3–5× higher at midday than at night.
  • Redundancy: at least N+1 replicas per region, multiple zones.
  • Headroom for bursts and prefill spikes.

Ways to raise utilisation:

  • Autoscale on queue depth or KV cache usage, within the limits of slow scale-up (minutes to load weights).
  • Fill the troughs with batch work: run evals, backfills, embeddings and batch-API-style jobs overnight on the same GPUs.
  • Consolidate models: serve several LoRA adapters on one base model instead of separate deployments (Module 9).
  • Multi-tenant platforms: pool many teams' traffic, so peaks average out.
  • Burst to APIs: size self-hosted capacity for the baseline and send spikes to a hosted API.

Hidden costs

  • Engineering and on-call: inference engine upgrades, driver and CUDA issues, capacity management, performance tuning. Often one or more full-time engineers, even for modest fleets.
  • Idle redundancy across zones and regions.
  • Storage and networking: model weights (hundreds of GB per model), fast startup (local NVMe caches), cross-zone traffic.
  • Observability and security: GPU metrics, request tracing, patching, isolation.
  • Evaluation and upgrades: every new open model release means an eval cycle and a migration.

Key takeaways

  • Self-hosted cost per million tokens = GPU $/hour ÷ (tokens per hour per GPU × utilisation).
  • Utilisation is everything. A fleet sized for peak and averaging 30% busy costs over 3× the ideal per token.
  • GPU pricing ranges widely, from on-demand and reserved to spot and owned hardware. Commitments cut price and add risk.
  • Hidden costs include engineering, idle redundancy, storage and networking, observability, and upgrades.

Go deeper

Finished reading? Mark it done to track your progress.