Design ChatGPT

In a Java service you scale by adding stateless replicas until QPS fits. An LLM replica is different: its throughput is set by how fast it can stream its weights and its KV cache out of GPU memory, so the number to derive is tokens per second per replica — and that depends on batch size, which depends on KV memory per request.

45 min round 9 computed numbers 4-part eval plan 8-point rubric
⚠️ Planning numbers, not measurements. Hardware and model facts are dated on the numbers sheet; traffic, prices and efficiency are labelled assumptions. Replace them with your own measurements (the vLLM load-test exercise) before quoting them.

The prompt

Design a ChatGPT-style chat service for 10 million daily users. State your numbers.

Attempt it first: the 45-minute round

Set the timer, answer out loud or on paper, then score yourself against the rubric before you read the model answer below.

  1. Framing (5 min) — Ask questions, fix the scale and the latency/quality/cost targets, state assumptions with numbers.
  2. Architecture (10 min) — Draw the boxes end to end: data in, the model call, storage, serving, feedback.
  3. Deep dive (15 min) — Pick the hardest part and do the arithmetic: tokens/s, KV memory, QPS → replicas, cost per 1,000 requests.
  4. Trade-offs (10 min) — Name what you would trade (batch vs latency, quality vs cost, build vs buy) and what breaks.
  5. Wrap-up (5 min) — Evaluation plan, monitoring, failure modes, and what you would do next.
45:00
Not started

Self-grade

Tick what you did. 0 of 8.

The model answer, step by step

Five steps, one tap each. The readout gives the step's answer and lists the numbers it uses; the same numbers are highlighted in the table underneath.

Design ChatGPT, one tap per step

👉 Predict first, then tap. Every number below is computed from the assumptions in the table.

1. Framing

5 min

2. Architecture

10 min

3. Deep dive

15 min

4. Trade-offs

10 min

5. Wrap-up

5 min

Tap a step above.
QuantityHow it is computedValue
Daily active usersassumption10,000,000
Requests per user per dayassumption8
Share of daily traffic in the peak hourassumption15%
Input tokens per request (history + prompt)assumption1,500
Output tokens per requestassumption400
GPUs per replica (tensor parallel), 70B at 1 byte per weightassumption4
Concurrent sequences per replica (batch)assumption128
Peak requests per seconddau × rpd × peak ÷ 36003,333
KV cache per request (fp16, 70B, in+out tokens)2 × 80 layers × 8 kv-heads × 128 × seq × 2 bytes623 MB
Most sequences that fit (90% of 4×80 GB, minus 70 GB weights)⌊(0.9 × 4 × 80 GB − 70 GB) ÷ KV per request⌋350
One decode step at that batch(70 GB + batch × KV) ÷ 4 GPUs ÷ 3.35 TB/s11.17 ms
Output tokens per second per replicabatch ÷ step11,458 tok/s
Output tokens per second needed at peakqps × output tokens1,333,333 tok/s
Replicas⌈needed ÷ per-replica⌉117
GPUsreplicas × 4468
Cost per 1,000 requests at $4/GPU-hourreplica $/hour ÷ (requests per hour per replica) × 1000$0.155
Takeaway. The headline figure — replicas — is 117 (⌈needed ÷ per-replica⌉). Say the assumption, show the formula, then give the number.

The evaluation plan

Go deeper on this site

Check yourself