Two counter-intuitive truths run every production AI system. A service gets dramatically slower long before it hits 100% busy — so you can never plan to run it full. And the cheapest request is the one you never send to the model at all. This page makes both visible, plus the fallback chain that keeps you up when the model doesn't answer.
Intuition says a server at 90% utilization is 90% as good as one at 50%. It is not. Requests arrive at
random, so they bunch up; when one arrives while the last is still being served, it waits in a queue.
The closer the arrival rate creeps to capacity, the longer those queues get — and near the top, average
response time doesn't rise, it explodes. This is the single most important curve in capacity planning,
and it's just W = 1 / (capacity − load).
Push the load toward capacity and watch response time go vertical:
The practical rules fall straight out of the curve. Plan for a headroom target — often around 60–70% utilization — because the region above it is where a small traffic spike becomes an outage. Autoscale on queue depth or latency, not just CPU, because the wall arrives before the CPU looks full. And put a queue with a timeout in front of the model so a backlog sheds load gracefully instead of melting down — which is exactly where the next section starts.
AI systems fail in ways ordinary services don't: variable latency, rate limits, the occasional refusal or malformed output. Reliability engineering's answer is reduced functionality over failure — a chain of fallbacks, each cheaper and more certain than the last, so the user gets a worse answer instead of an error page. Step down the chain as the primary path fails:
Each rung trades quality for certainty: the flagship model is best but slowest and least reliable; a smaller model is faster and cheaper; a cached answer is instant; a templated response always works. The art is timeouts and retries with backoff that fail fast enough to try the next rung within your latency budget, plus idempotency so a retry can't double-charge or double-send.
At volume, model calls are the dominant line item, and the highest-leverage optimisation is a semantic cache: when a new request means the same as one you've answered before — not the same string, the same meaning — you return the stored answer for the price of an embedding lookup. Even a modest hit rate carves a large slice off the bill. Model your monthly cost and turn up the cache:
W = 1/(μ−λ)) — real systems add concurrency, batching and variable service times, but the
shape and the "wall" are exactly right · the cost model counts only the cached fraction as saved and
charges an embedding lookup on every request; it omits egress, vector-DB and retry costs · a semantic cache
trades freshness for cost — too loose a similarity threshold returns wrong answers, so hit rate and staleness
are a real tension · fallbacks must be tested by actually forcing the primary to fail, not assumed.Second opinion (the topic is taught here — these corroborate): System Design Primer — scalability · FinOps Foundation · Ray Serve — scaling.