A Java request/response service can take a second. A voice agent has a conversation rhythm: past about a second of silence it feels broken. The design is a latency budget split across streaming stages, and the fleet is sized by concurrent calls, not requests per second.
45 min round7 computed numbers4-part eval plan8-point rubric
⚠️ Planning numbers, not measurements. Hardware and model facts
are dated on the numbers sheet; traffic, prices and efficiency are labelled
assumptions. Replace them with your own measurements (the vLLM load-test exercise) before quoting them.
The prompt
Design a phone voice agent that talks with 2,000 callers at the same time.
Attempt it first: the 45-minute round
Set the timer, answer out loud or on paper, then score yourself against the rubric before you read the model answer below.
Framing (5 min) — Ask questions, fix the scale and the latency/quality/cost targets, state assumptions with numbers.
Architecture (10 min) — Draw the boxes end to end: data in, the model call, storage, serving, feedback.
Deep dive (15 min) — Pick the hardest part and do the arithmetic: tokens/s, KV memory, QPS → replicas, cost per 1,000 requests.
Trade-offs (10 min) — Name what you would trade (batch vs latency, quality vs cost, build vs buy) and what breaks.
Wrap-up (5 min) — Evaluation plan, monitoring, failure modes, and what you would do next.
45:00
Not started
Self-grade
Tick what you did. 0 of 8.
The model answer, step by step
Five steps, one tap each. The readout gives the step's answer and lists the numbers it uses; the same numbers are highlighted in the table underneath.
Real-time voice agent, one tap per step
👉 Predict first, then tap. Every number below is computed from the assumptions in the table.
1. Framing
5 min
2. Architecture
10 min
3. Deep dive
15 min
4. Trade-offs
10 min
5. Wrap-up
5 min
Tap a step above.
Quantity
How it is computed
Value
Peak concurrent calls
assumption
2,000
Seconds between agent turns per call
assumption
10
LLM input tokens per turn (history)
assumption
1,200
LLM output tokens per turn
assumption
60
Decode batch
assumption
32
Utilisation
assumption
60%
End-of-speech detection, ms
assumption
200
Speech-to-text finalisation, ms
assumption
150
LLM first token, ms
assumption
250
Text-to-speech first audio, ms
assumption
150
Streams per GPU for speech models (assumed)
assumption
40
Silence before the caller hears a reply
sum of the four stages
750 ms
LLM turns per second
calls ÷ gap
200
GPU-seconds per LLM turn (8B)
prefill + decode
0.0599 s
LLM GPUs
turns × GPU-s ÷ utilisation
20
GPUs each for STT and TTS
⌈calls ÷ streams per GPU⌉
50
Total GPUs
LLM + STT + TTS
120
GPU cost per call-minute at $4/GPU-hour
GPUs × $4 ÷ (calls × 60)
$0.0040
Takeaway. The headline figure — total gpus — is 120
(LLM + STT + TTS). Say the assumption, show the formula, then give the number.
The evaluation plan
Speech-to-text word error rate on 500 real calls with accents and noise.
End-to-end scripted calls: task completion and turn-latency percentiles.
Barge-in test set: interruptions should stop the agent within 300 ms.
Human rating of naturalness and of unsafe or off-policy statements.