A Java system-design answer sizes servers from requests per second. A GenAI answer has to size them from tokens per second, KV-cache memory and cost per thousand requests. Twelve designs, each worked with real arithmetic, a 45-minute timer, a self-grade rubric and a model answer to compare with.
| # | Step | Minutes | What you do |
|---|---|---|---|
| 1 | Framing | 5 min | Ask questions, fix the scale and the latency/quality/cost targets, state assumptions with numbers. |
| 2 | Architecture | 10 min | Draw the boxes end to end: data in, the model call, storage, serving, feedback. |
| 3 | Deep dive | 15 min | Pick the hardest part and do the arithmetic: tokens/s, KV memory, QPS → replicas, cost per 1,000 requests. |
| 4 | Trade-offs | 10 min | Name what you would trade (batch vs latency, quality vs cost, build vs buy) and what breaks. |
| 5 | Wrap-up | 5 min | Evaluation plan, monitoring, failure modes, and what you would do next. |
The same five steps are used by every design page. The wider framework and templates are on ML & GenAI system design and Interview practice; the six system-design write-ups on The Path use the same numbers sheet.
| Quantity | Formula |
|---|---|
| KV cache | 2 × layers × KV heads × head dim × tokens × batch × bytes per value |
| Decode step time | (weight bytes + batch × KV bytes) ÷ GPUs ÷ memory bandwidth — decode is memory-bound |
| Tokens/s per replica | batch ÷ step time |
| Prefill GPU-seconds | 2 × parameters × input tokens ÷ achieved FLOP/s |
| Replicas | ⌈needed tokens/s ÷ tokens/s per replica⌉ (or ⌈QPS ÷ per-replica QPS⌉) |
| Concurrency | arrival rate × duration (Little's law) |
| Cost per 1,000 requests | GPU-seconds per request × 1000 ÷ 3600 × $ per GPU-hour |
Facts read on 2026-09-30: H100 SXM memory bandwidth 3.35 TB/s; Llama 3 8B has 32 layers and 8 KV heads with head dimension 128, and 70B has 80 layers, 8 KV heads and head dimension 128. Assumptions (not measurements): $4 per GPU-hour, 400 TFLOP/s achieved prefill, traffic shapes and API prices, each labelled on its design page. Replace them with your own numbers from the vLLM load-test exercise before quoting.
Open a design, read only its prompt, start the 45-minute timer and answer. Then tick the rubric honestly and compare with the model answer. Twelve designs is twelve rounds — do one a day, and redo any where you scored under five.