"Should we call an API or host our own model?" is the most consequential — and most badly argued — decision on an AI project. It usually gets settled by whoever is loudest, then re-litigated every quarter. This page turns it into something you can show: a cost curve with a crossover point, and a weighted matrix that makes your priorities visible instead of implicit.
A managed API charges per token — near-zero to start, and it scales linearly forever. Self-hosting costs a fixed amount per month (GPU, or the engineer keeping it alive) and then serves additional requests almost free. Two different shapes: a line through the origin, and a line starting high but nearly flat. Two such lines cross exactly once.
Below that crossover, self-hosting is a waste of money. Above it, the API is. The only question that matters is which side of the crossover your actual volume sits on — and whether it will still be there in a year. Turn the dials:
Two honest warnings about that number. The fixed cost is not the GPU bill — it's the GPU plus the fraction of an engineer who patches, monitors and upgrades it, and that person is usually more expensive than the hardware. And the crossover moves: API prices have fallen steeply year over year, so a decision that was marginal at today's prices can invert before your amortisation period ends. Decide with a range, not a point.
Teams argue about build-vs-buy because they're weighting different things silently — one person optimising for cost, another for data privacy, a third for shipping speed. A weighted decision matrix forces the weights into the open, and the moment they're explicit the argument usually resolves itself.
Set the weights to what your organisation actually cares about and watch the winner change. This is the artefact you paste into an Architecture Decision Record — not the answer, but the reasoning:
Over-engineering is building a Kubernetes-and-vLLM platform to serve 200 requests a day — you have bought yourself an operational burden with no cost saving, and shipped two quarters late. Under-resourcing is the mirror image: choosing self-hosting for the cost curve, then discovering nobody owns the pager. The decision matrix catches both, because "team capability to operate" is a criterion with a weight rather than an afterthought.
The compromise most mature teams land on is neither column: start on an API behind your own gateway interface, so the provider is swappable. You buy speed now and keep the option to build later, once volume tells you where the crossover actually is. That's the "loose coupling" pattern — the decision you don't have to make yet is worth real money.
Second opinion (read these after you've worked the models above — the topic is taught here, these are for corroboration): ADR collection (Nygard format) · FinOps Foundation — what is FinOps · AWS Generative AI Lens.