In a Java service you scale by adding stateless replicas until QPS fits. An LLM replica is different: its throughput is set by how fast it can stream its weights and its KV cache out of GPU memory, so the number to derive is tokens per second per replica — and that depends on batch size, which depends on KV memory per request.
Design a ChatGPT-style chat service for 10 million daily users. State your numbers.
Set the timer, answer out loud or on paper, then score yourself against the rubric before you read the model answer below.
Five steps, one tap each. The readout gives the step's answer and lists the numbers it uses; the same numbers are highlighted in the table underneath.
5 min
10 min
15 min
10 min
5 min
| Quantity | How it is computed | Value |
|---|---|---|
| Daily active users | assumption | 10,000,000 |
| Requests per user per day | assumption | 8 |
| Share of daily traffic in the peak hour | assumption | 15% |
| Input tokens per request (history + prompt) | assumption | 1,500 |
| Output tokens per request | assumption | 400 |
| GPUs per replica (tensor parallel), 70B at 1 byte per weight | assumption | 4 |
| Concurrent sequences per replica (batch) | assumption | 128 |
| Peak requests per second | dau × rpd × peak ÷ 3600 | 3,333 |
| KV cache per request (fp16, 70B, in+out tokens) | 2 × 80 layers × 8 kv-heads × 128 × seq × 2 bytes | 623 MB |
| Most sequences that fit (90% of 4×80 GB, minus 70 GB weights) | ⌊(0.9 × 4 × 80 GB − 70 GB) ÷ KV per request⌋ | 350 |
| One decode step at that batch | (70 GB + batch × KV) ÷ 4 GPUs ÷ 3.35 TB/s | 11.17 ms |
| Output tokens per second per replica | batch ÷ step | 11,458 tok/s |
| Output tokens per second needed at peak | qps × output tokens | 1,333,333 tok/s |
| Replicas | ⌈needed ÷ per-replica⌉ | 117 |
| GPUs | replicas × 4 | 468 |
| Cost per 1,000 requests at $4/GPU-hour | replica $/hour ÷ (requests per hour per replica) × 1000 | $0.155 |