Build vs buy

"Should we call an API or host our own model?" is the most consequential — and most badly argued — decision on an AI project. It usually gets settled by whoever is loudest, then re-litigated every quarter. This page turns it into something you can show: a cost curve with a crossover point, and a weighted matrix that makes your priorities visible instead of implicit.

build vs buydecision matrix ADRcost at volumevendor lock-in

First: the crossover is real, and you can compute it

A managed API charges per token — near-zero to start, and it scales linearly forever. Self-hosting costs a fixed amount per month (GPU, or the engineer keeping it alive) and then serves additional requests almost free. Two different shapes: a line through the origin, and a line starting high but nearly flat. Two such lines cross exactly once.

Below that crossover, self-hosting is a waste of money. Above it, the API is. The only question that matters is which side of the crossover your actual volume sits on — and whether it will still be there in a year. Turn the dials:

Where does your volume cross over?

prompt + output; RAG context makes this big fast
GPU + the engineer time to run it — the part teams forget

Two honest warnings about that number. The fixed cost is not the GPU bill — it's the GPU plus the fraction of an engineer who patches, monitors and upgrades it, and that person is usually more expensive than the hardware. And the crossover moves: API prices have fallen steeply year over year, so a decision that was marginal at today's prices can invert before your amortisation period ends. Decide with a range, not a point.

Cost is one criterion. Make the rest visible.

Teams argue about build-vs-buy because they're weighting different things silently — one person optimising for cost, another for data privacy, a third for shipping speed. A weighted decision matrix forces the weights into the open, and the moment they're explicit the argument usually resolves itself.

Set the weights to what your organisation actually cares about and watch the winner change. This is the artefact you paste into an Architecture Decision Record — not the answer, but the reasoning:

Weighted decision matrix — set your priorities, see the winner

be honest — this is the criterion teams flatter themselves on

The two failure modes

Over-engineering is building a Kubernetes-and-vLLM platform to serve 200 requests a day — you have bought yourself an operational burden with no cost saving, and shipped two quarters late. Under-resourcing is the mirror image: choosing self-hosting for the cost curve, then discovering nobody owns the pager. The decision matrix catches both, because "team capability to operate" is a criterion with a weight rather than an afterthought.

The compromise most mature teams land on is neither column: start on an API behind your own gateway interface, so the provider is swappable. You buy speed now and keep the option to build later, once volume tells you where the crossover actually is. That's the "loose coupling" pattern — the decision you don't have to make yet is worth real money.

⚠️ Traps & honesty: the cost model here is deliberately simple — linear API pricing vs a flat self-host cost. It omits egress, vector-DB and retrieval cost, cache hit rates, reserved- instance discounts and batch pricing, all of which move the crossover · the numbers are yours to supply: the page computes the crossover from the dials, it does not claim your prices · matrix scores are a judgement, not a measurement — their value is that they are written down and can be argued with · a decision matrix cannot rescue a badly chosen criteria list.
Takeaways: API cost is linear in volume, self-hosting is fixed-plus-flat, so there is exactly one crossover — compute it before arguing · include engineer time in the fixed cost, not just the GPU · make the weights explicit in a matrix so people are arguing about priorities rather than conclusions · record the choice, the alternatives and the reasoning in an ADR so it can be revisited when prices move · the default that preserves optionality is an API behind a swappable gateway. Next: the MLOps lifecycle is where the system you chose actually runs.

Second opinion (read these after you've worked the models above — the topic is taught here, these are for corroboration): ADR collection (Nygard format) · FinOps Foundation — what is FinOps · AWS Generative AI Lens.