API vs Self-Hosted Models
The trade-offs between hosted model APIs, managed cloud endpoints and running open-weight models yourself, and a decision checklist for interviews.
Lesson 6 of 6 9 min
Three ways to get a model
Provider API
OpenAI, Anthropic, Google…
Managed cloud endpoint
Bedrock, Vertex AI, Azure; dedicated capacity
Self-hosted open weights
vLLM/SGLang on your GPUs
- Provider API: pay per token, get the newest frontier models, and the provider runs everything.
- Managed cloud endpoint: the same or similar models served through your cloud provider. Traffic stays in your cloud account and region, billing is consolidated, and there are options for provisioned throughput (reserved capacity at a fixed hourly price).
- Self-hosted: download open weights (Llama, Qwen, Mistral, Gemma, gpt-oss, DeepSeek and others) and serve them with an inference engine on GPUs you rent or own.
Comparison
| Factor | Provider API | Self-hosted |
|---|---|---|
| Best available quality | Frontier models | Open models, strong but often behind the very top |
| Time to production | Hours | Weeks |
| Cost at low or spiky volume | Low: pay per use | High: idle GPUs |
| Cost at high, steady volume | High | Often several times cheaper per token |
| Scaling | Instant, within rate limits | Minutes to hours; GPU availability can be the limit |
| Latency control | Shared infrastructure, variable | Tunable (batch size, hardware, placement) |
| Data control | Contracts, retention settings, regions | Complete. Can run air-gapped |
| Customisation | Prompting, some hosted fine-tuning | Full fine-tuning, LoRA adapters, custom decoding |
| Ops burden | None | GPUs, drivers, engine upgrades, autoscaling, on-call |
The break-even question
Self-hosting is a fixed cost and APIs are a variable cost. Roughly:
Self-hosting pays off when (monthly API spend) > (GPU fleet cost at your peak) + (engineering and ops cost)
Two things make this harder than it looks:
- You pay for peak, not average. If traffic peaks at 3× the average, your GPUs average about 33% utilisation, which triples the effective cost per token.
- People cost money. Running a reliable GPU inference platform takes skilled engineers. At a small scale that cost dwarfs the GPU bill.
Module 2's capacity planning lesson works through a case where self-hosting is ≈ 8× cheaper at 10M requests a day, and Module 8 covers break-even in depth.
When self-hosting is right regardless of cost
- Data can't leave your boundary: regulated industries, government, air-gapped environments, strict data-residency rules.
- Latency or placement needs: on-premise, edge devices, or co-location with other systems.
- Deep customisation: heavy fine-tuning, custom architectures, many LoRA adapters, custom sampling.
- Very high volume of simple tasks: a fine-tuned small model on your own GPUs can be orders of magnitude cheaper than a general API model.
The hybrid default
Many companies end up with a mix:
- A gateway in front of everything, with one internal API and provider-specific adapters behind it.
- Hosted frontier models for hard, low-volume tasks.
- Self-hosted or provisioned small and mid models for high-volume, predictable workloads.
- Failover between providers or regions for resilience.
Key takeaways
- There are three options: a provider API, a managed endpoint in your cloud, or self-hosting open weights on your own GPUs.
- APIs win on quality, elasticity and zero ops. Self-hosting wins on cost at high steady volume, data control and customisation.
- The self-hosting break-even depends on utilisation, because idle GPUs cost the same as busy ones.
- A gateway that abstracts the provider lets you mix options and switch later.
Go deeper
Finished reading? Mark it done to track your progress.