Self-Hosting Costs: GPU-Hours
What self-hosting really costs, including GPU pricing models, utilisation, cost per million tokens from throughput, the hidden costs, and how to raise utilisation.
Lesson 3 of 5 9 min
From GPU-hours to cost per token
Cost per 1M output tokens = (GPU $/hour × GPUs per replica) ÷ (replica tokens/s × 3,600 × utilisation) × 1,000,000
Example: Llama 3.1 70B in FP8 on 2 × H100 at $2.50 per GPU-hour, with ≈ 2,500 output tokens/s per replica (from the capacity planning lesson):
- Replica cost: $5 an hour
- At 100% utilisation: 2,500 × 3,600 = 9M tokens an hour, so ≈ $0.56 per million output tokens
- At 40% average utilisation: ≈ $1.39 per million
- At 15% utilisation (low traffic with redundancy): ≈ $3.70 per million
Compare that with API output prices for similar-quality models to see which side of the line you're on. Input (prefill) tokens are much cheaper to process per token than output, so for a like-for-like comparison, count both.
GPU pricing models
| Model | Typical discount vs on-demand | Trade-off |
|---|---|---|
| On-demand cloud | — | Flexible, but availability of top GPUs isn't guaranteed |
| Reserved / committed (1–3 years) | Large, often 30–60% | Locked in as prices fall and hardware improves |
| Spot / preemptible | Very large | Can be reclaimed at short notice. Suits batch, not serving |
| Specialised GPU clouds | Often cheaper than hyperscalers | Fewer surrounding services, varying reliability |
| Owned hardware | Lowest at high utilisation over years | Capex, data-centre power and cooling, operations, refresh cycles |
Utilisation: the real driver
You provision for peak plus redundancy, but pay for every hour:
- Diurnal patterns: traffic may be 3–5× higher at midday than at night.
- Redundancy: at least N+1 replicas per region, multiple zones.
- Headroom for bursts and prefill spikes.
Ways to raise utilisation:
- Autoscale on queue depth or KV cache usage, within the limits of slow scale-up (minutes to load weights).
- Fill the troughs with batch work: run evals, backfills, embeddings and batch-API-style jobs overnight on the same GPUs.
- Consolidate models: serve several LoRA adapters on one base model instead of separate deployments (Module 9).
- Multi-tenant platforms: pool many teams' traffic, so peaks average out.
- Burst to APIs: size self-hosted capacity for the baseline and send spikes to a hosted API.
Hidden costs
- Engineering and on-call: inference engine upgrades, driver and CUDA issues, capacity management, performance tuning. Often one or more full-time engineers, even for modest fleets.
- Idle redundancy across zones and regions.
- Storage and networking: model weights (hundreds of GB per model), fast startup (local NVMe caches), cross-zone traffic.
- Observability and security: GPU metrics, request tracing, patching, isolation.
- Evaluation and upgrades: every new open model release means an eval cycle and a migration.
Key takeaways
- Self-hosted cost per million tokens = GPU $/hour ÷ (tokens per hour per GPU × utilisation).
- Utilisation is everything. A fleet sized for peak and averaging 30% busy costs over 3× the ideal per token.
- GPU pricing ranges widely, from on-demand and reserved to spot and owned hardware. Commitments cut price and add risk.
- Hidden costs include engineering, idle redundancy, storage and networking, observability, and upgrades.
Go deeper
Finished reading? Mark it done to track your progress.