Cost, scale & reliability

Two counter-intuitive truths run every production AI system. A service gets dramatically slower long before it hits 100% busy — so you can never plan to run it full. And the cheapest request is the one you never send to the model at all. This page makes both visible, plus the fallback chain that keeps you up when the model doesn't answer.

queueingutilizationgraceful degradation semantic cacheFinOps

The latency wall: you cannot run at 100%

Intuition says a server at 90% utilization is 90% as good as one at 50%. It is not. Requests arrive at random, so they bunch up; when one arrives while the last is still being served, it waits in a queue. The closer the arrival rate creeps to capacity, the longer those queues get — and near the top, average response time doesn't rise, it explodes. This is the single most important curve in capacity planning, and it's just W = 1 / (capacity − load).

Push the load toward capacity and watch response time go vertical:

Response time vs load

watch what happens as this approaches capacity

The practical rules fall straight out of the curve. Plan for a headroom target — often around 60–70% utilization — because the region above it is where a small traffic spike becomes an outage. Autoscale on queue depth or latency, not just CPU, because the wall arrives before the CPU looks full. And put a queue with a timeout in front of the model so a backlog sheds load gracefully instead of melting down — which is exactly where the next section starts.

Graceful degradation: always have a worse answer ready

AI systems fail in ways ordinary services don't: variable latency, rate limits, the occasional refusal or malformed output. Reliability engineering's answer is reduced functionality over failure — a chain of fallbacks, each cheaper and more certain than the last, so the user gets a worse answer instead of an error page. Step down the chain as the primary path fails:

The fallback chain — ▶ play it

Each rung trades quality for certainty: the flagship model is best but slowest and least reliable; a smaller model is faster and cheaper; a cached answer is instant; a templated response always works. The art is timeouts and retries with backoff that fail fast enough to try the next rung within your latency budget, plus idempotency so a retry can't double-charge or double-send.

FinOps: the cheapest token is the one you don't spend

At volume, model calls are the dominant line item, and the highest-leverage optimisation is a semantic cache: when a new request means the same as one you've answered before — not the same string, the same meaning — you return the stored answer for the price of an embedding lookup. Even a modest hit rate carves a large slice off the bill. Model your monthly cost and turn up the cache:

Monthly cost, with a semantic cache

fraction of requests a cached answer can serve

⚠️ Traps & honesty: the latency curve is the standard M/M/1 queue (W = 1/(μ−λ)) — real systems add concurrency, batching and variable service times, but the shape and the "wall" are exactly right · the cost model counts only the cached fraction as saved and charges an embedding lookup on every request; it omits egress, vector-DB and retry costs · a semantic cache trades freshness for cost — too loose a similarity threshold returns wrong answers, so hit rate and staleness are a real tension · fallbacks must be tested by actually forcing the primary to fail, not assumed.
Takeaways: response time explodes as load nears capacity — plan for headroom (~60–70%), never 100%, and autoscale on latency/queue depth · put a queue + timeout in front so overload sheds gracefully · build a fallback chain (flagship → small model → cache → template) that returns a worse answer instead of an error, with fast timeouts and idempotent retries · a semantic cache is the biggest FinOps lever — even 30% hit rate is 30% off the model bill. Next: AI governance puts guardrails around the system you've now scaled.

Second opinion (the topic is taught here — these corroborate): System Design Primer — scalability · FinOps Foundation · Ray Serve — scaling.