Build vs Buy
A structured way to decide between hosted APIs and self-hosting, covering the cost break-even, quality, control, risk and team capability, plus the hybrid strategies most companies land on.
The shape of the two curves
Build vs buy break-even
Hosted API cost grows linearly with traffic. Self-hosting is a fixed floor (redundant replicas plus engineering) that rises in steps as you add replicas.
From a benchmark or the roofline estimate
People to run the platform
Below break-even, self-hosting pays for idle GPUs and a platform team. Above it, API list prices outgrow a fleet sized for peak. Prompt caching, batch discounts and committed-use GPU pricing all move the line, so plug in your real numbers.
- API ("buy"): a straight line. Twice the traffic costs twice as much.
- Self-hosted ("build"): starts at a floor made of at least two replicas for redundancy plus the team to run them, then rises in steps as replicas are added.
Below the crossover you pay for idle GPUs and a platform team. Above it, per-token API prices outgrow a fleet that grows in chunks. With the defaults above (a 70B-class model on 2 × H100 against a mid-tier API), the crossover lands around a couple of hundred thousand requests a day. Change the inputs and it moves by orders of magnitude.
A decision checklist
| Factor | Favours API | Favours self-hosting |
|---|---|---|
| Volume | Low, spiky, or uncertain | High and steady |
| Quality needed | Frontier-level reasoning | Open models pass your evals |
| Data constraints | Contractual controls are sufficient | Data can't leave your boundary |
| Latency | Standard | Special placement, on-prem or edge, tight tail control |
| Customisation | Prompting, light hosted fine-tuning | Heavy fine-tuning, many adapters, custom decoding |
| Team | No GPU or inference expertise | An existing ML platform team |
| Speed | Ship this quarter | Can invest months |
| Risk tolerance | Prefer vendor SLAs | Prefer control over dependencies |
Hybrid strategies
- Tiered by task: self-host a fine-tuned small model for high-volume classification and extraction, and use a frontier API for complex reasoning.
- Baseline plus burst: self-host (or provision) capacity for steady load, and overflow to a pay-as-you-go API at peaks.
- Primary plus fallback: API primary with a self-hosted fallback for outages, or the reverse.
- Graduation: start every feature on an API, and move a workload to self-hosting only when it is proven, stable, high-volume and an open model passes evals.
All of these depend on a gateway with a provider-neutral API (Module 6), so moving a workload is a routing change, not a rewrite.
Risks on each side
API risks: price changes, model deprecations (forced migrations), rate limits at critical moments, outages, data-policy changes, and vendor lock-in through provider-specific features.
Self-hosting risks: GPU availability, operational incidents you must fix yourself, falling behind frontier quality, hardware commitments that age badly, and key-person dependency on a small platform team.
Mitigations: pin model versions, keep evals ready to validate alternatives, abstract providers behind the gateway, and avoid multi-year hardware commitments unless utilisation is proven.
Key takeaways
- API cost scales linearly with traffic. Self-hosting is a step function above a fixed floor of redundant replicas plus engineering.
- The break-even point depends heavily on utilisation, model quality needs, and engineering cost. Compute it, don't guess.
- Non-cost factors often decide. Consider data control, latency, customisation, model quality, and your team's ability to operate GPUs.
- Most companies end up hybrid. They self-host high-volume, stable workloads and use APIs for frontier quality, spikes and experimentation.