GenAI system-design round

A Java system-design answer sizes servers from requests per second. A GenAI answer has to size them from tokens per second, KV-cache memory and cost per thousand requests. Twelve designs, each worked with real arithmetic, a 45-minute timer, a self-grade rubric and a model answer to compare with.

12 worked designs45-minute timerarithmetic script-checked

The five steps

#StepMinutesWhat you do
1Framing5 minAsk questions, fix the scale and the latency/quality/cost targets, state assumptions with numbers.
2Architecture10 minDraw the boxes end to end: data in, the model call, storage, serving, feedback.
3Deep dive15 minPick the hardest part and do the arithmetic: tokens/s, KV memory, QPS → replicas, cost per 1,000 requests.
4Trade-offs10 minName what you would trade (batch vs latency, quality vs cost, build vs buy) and what breaks.
5Wrap-up5 minEvaluation plan, monitoring, failure modes, and what you would do next.

The same five steps are used by every design page. The wider framework and templates are on ML & GenAI system design and Interview practice; the six system-design write-ups on The Path use the same numbers sheet.

The numbers sheet

QuantityFormula
KV cache2 × layers × KV heads × head dim × tokens × batch × bytes per value
Decode step time(weight bytes + batch × KV bytes) ÷ GPUs ÷ memory bandwidth — decode is memory-bound
Tokens/s per replicabatch ÷ step time
Prefill GPU-seconds2 × parameters × input tokens ÷ achieved FLOP/s
Replicas⌈needed tokens/s ÷ tokens/s per replica⌉ (or ⌈QPS ÷ per-replica QPS⌉)
Concurrencyarrival rate × duration (Little's law)
Cost per 1,000 requestsGPU-seconds per request × 1000 ÷ 3600 × $ per GPU-hour

Facts read on 2026-09-30: H100 SXM memory bandwidth 3.35 TB/s; Llama 3 8B has 32 layers and 8 KV heads with head dimension 128, and 70B has 80 layers, 8 KV heads and head dimension 128. Assumptions (not measurements): $4 per GPU-hour, 400 TFLOP/s achieved prefill, traffic shapes and API prices, each labelled on its design page. Replace them with your own numbers from the vLLM load-test exercise before quoting.

Pick a design

1. Design ChatGPTa consumer chat product on a 70B model. 2. Enterprise search / RAG at scalepermission-aware Q&A over 20M chunks. 3. Customer-support agentan agent with tools, priced per conversation. 4. Document-processing pipeline2M scanned pages a day into structured JSON. 5. Coding assistantautocomplete under 300 ms plus chat. 6. LLM gateway with provider failoverone front door over two model providers. 7. Eval platformregression testing for prompts and models. 8. Multimodal searchtext-to-image search over 100M photos. 9. LLM-powered recommendationsan LLM reranker inside a 150 ms budget. 10. On-device assistant, privacy-firsta 3B model on the phone, cloud only as fallback. 11. Content moderationa cascade from cheap classifier to LLM to humans. 12. Real-time voice agentan 800 ms turn budget across STT, LLM and TTS.

Timed practice

Open a design, read only its prompt, start the 45-minute timer and answer. Then tick the rubric honestly and compare with the model answer. Twelve designs is twelve rounds — do one a day, and redo any where you scored under five.

Sources

  1. Meta — The Llama 3 Herd of Models, Table 3 (layers, KV heads, model dimension) — retrieved 2026-09-30
  2. NVIDIA — H100 Tensor Core GPU (3.35 TB/s memory bandwidth, SXM) — retrieved 2026-09-30
  3. IGotAnOffer — Generative AI system design interview (the five-step shape) — as cited on the site's Careers page, 2026-09-30