A Java shop puts a service mesh or API gateway in front of internal services. An LLM gateway does that plus three things a mesh does not: it counts tokens as well as requests, it holds long streaming connections open, and it decides which vendor's model answers.
45 min round6 computed numbers4-part eval plan8-point rubric
⚠️ Planning numbers, not measurements. Hardware and model facts
are dated on the numbers sheet; traffic, prices and efficiency are labelled
assumptions. Replace them with your own measurements (the vLLM load-test exercise) before quoting them.
The prompt
Design an internal gateway that gives every team one API over multiple LLM providers, with rate limits and failover.
Attempt it first: the 45-minute round
Set the timer, answer out loud or on paper, then score yourself against the rubric before you read the model answer below.
Framing (5 min) — Ask questions, fix the scale and the latency/quality/cost targets, state assumptions with numbers.
Architecture (10 min) — Draw the boxes end to end: data in, the model call, storage, serving, feedback.
Deep dive (15 min) — Pick the hardest part and do the arithmetic: tokens/s, KV memory, QPS → replicas, cost per 1,000 requests.
Trade-offs (10 min) — Name what you would trade (batch vs latency, quality vs cost, build vs buy) and what breaks.
Wrap-up (5 min) — Evaluation plan, monitoring, failure modes, and what you would do next.
45:00
Not started
Self-grade
Tick what you did. 0 of 8.
The model answer, step by step
Five steps, one tap each. The readout gives the step's answer and lists the numbers it uses; the same numbers are highlighted in the table underneath.
LLM gateway with provider failover, one tap per step
👉 Predict first, then tap. Every number below is computed from the assumptions in the table.
Tokens per minute one provider key allows (assumed)
assumption
4,000,000
Concurrent open streams (Little's law)
rps × duration
2,400
Nodes to hold them
⌈streams ÷ per-node⌉
3
Nodes so losing one AZ still fits
⌈need × azs ÷ (azs − 1)⌉
5
Requests failing on both providers
pFail × pFail
0.04%
Tokens per minute at peak
rps × 60 × tokens
27,000,000
Provider keys/accounts needed
⌈tokens/min ÷ per-key limit⌉
7
Takeaway. The headline figure — nodes so losing one az still fits — is 5
(⌈need × azs ÷ (azs − 1)⌉). Say the assumption, show the formula, then give the number.
The evaluation plan
Contract tests: the same 100 requests through each provider adapter must yield equivalent structure and error mapping.
Chaos drills: kill a provider, verify traffic shifts within the target seconds and the error rate stays under the SLO.
Quality parity: sample outputs from fallback models scored against the primary on the golden set, so a silent quality drop is visible.
Load test to 1.5× peak with limits on, verifying 429s carry a retry-after.