GenAI System Design
1. LLM Fundamentals for System Designers

API vs Self-Hosted Models

The trade-offs between hosted model APIs, managed cloud endpoints and running open-weight models yourself, and a decision checklist for interviews.

Lesson 6 of 6 9 min

Three ways to get a model

Provider API
OpenAI, Anthropic, Google…
Managed cloud endpoint
Bedrock, Vertex AI, Azure; dedicated capacity
Self-hosted open weights
vLLM/SGLang on your GPUs
Left to right: less operational work → more control and, at scale, lower unit cost.
  • Provider API: pay per token, get the newest frontier models, and the provider runs everything.
  • Managed cloud endpoint: the same or similar models served through your cloud provider. Traffic stays in your cloud account and region, billing is consolidated, and there are options for provisioned throughput (reserved capacity at a fixed hourly price).
  • Self-hosted: download open weights (Llama, Qwen, Mistral, Gemma, gpt-oss, DeepSeek and others) and serve them with an inference engine on GPUs you rent or own.

Comparison

FactorProvider APISelf-hosted
Best available qualityFrontier modelsOpen models, strong but often behind the very top
Time to productionHoursWeeks
Cost at low or spiky volumeLow: pay per useHigh: idle GPUs
Cost at high, steady volumeHighOften several times cheaper per token
ScalingInstant, within rate limitsMinutes to hours; GPU availability can be the limit
Latency controlShared infrastructure, variableTunable (batch size, hardware, placement)
Data controlContracts, retention settings, regionsComplete. Can run air-gapped
CustomisationPrompting, some hosted fine-tuningFull fine-tuning, LoRA adapters, custom decoding
Ops burdenNoneGPUs, drivers, engine upgrades, autoscaling, on-call

The break-even question

Self-hosting is a fixed cost and APIs are a variable cost. Roughly:

Self-hosting pays off when (monthly API spend) > (GPU fleet cost at your peak) + (engineering and ops cost)

Two things make this harder than it looks:

  1. You pay for peak, not average. If traffic peaks at 3× the average, your GPUs average about 33% utilisation, which triples the effective cost per token.
  2. People cost money. Running a reliable GPU inference platform takes skilled engineers. At a small scale that cost dwarfs the GPU bill.

Module 2's capacity planning lesson works through a case where self-hosting is ≈ 8× cheaper at 10M requests a day, and Module 8 covers break-even in depth.

When self-hosting is right regardless of cost

  • Data can't leave your boundary: regulated industries, government, air-gapped environments, strict data-residency rules.
  • Latency or placement needs: on-premise, edge devices, or co-location with other systems.
  • Deep customisation: heavy fine-tuning, custom architectures, many LoRA adapters, custom sampling.
  • Very high volume of simple tasks: a fine-tuned small model on your own GPUs can be orders of magnitude cheaper than a general API model.

The hybrid default

Many companies end up with a mix:

  • A gateway in front of everything, with one internal API and provider-specific adapters behind it.
  • Hosted frontier models for hard, low-volume tasks.
  • Self-hosted or provisioned small and mid models for high-volume, predictable workloads.
  • Failover between providers or regions for resilience.

Key takeaways

  • There are three options: a provider API, a managed endpoint in your cloud, or self-hosting open weights on your own GPUs.
  • APIs win on quality, elasticity and zero ops. Self-hosting wins on cost at high steady volume, data control and customisation.
  • The self-hosting break-even depends on utilisation, because idle GPUs cost the same as busy ones.
  • A gateway that abstracts the provider lets you mix options and switch later.

Go deeper

Finished reading? Mark it done to track your progress.