In a Java service one request is one call. An agent turns one conversation into a loop of model calls, each carrying the growing history, so cost and latency are multiplied by the number of steps — and the tool calls have side effects a retry can repeat.
45 min round5 computed numbers4-part eval plan8-point rubric
⚠️ Planning numbers, not measurements. Hardware and model facts
are dated on the numbers sheet; traffic, prices and efficiency are labelled
assumptions. Replace them with your own measurements (the vLLM load-test exercise) before quoting them.
The prompt
Design an agent that resolves customer-support chats using tools (order lookup, refunds, knowledge base).
Attempt it first: the 45-minute round
Set the timer, answer out loud or on paper, then score yourself against the rubric before you read the model answer below.
Framing (5 min) — Ask questions, fix the scale and the latency/quality/cost targets, state assumptions with numbers.
Architecture (10 min) — Draw the boxes end to end: data in, the model call, storage, serving, feedback.
Deep dive (15 min) — Pick the hardest part and do the arithmetic: tokens/s, KV memory, QPS → replicas, cost per 1,000 requests.
Trade-offs (10 min) — Name what you would trade (batch vs latency, quality vs cost, build vs buy) and what breaks.
Wrap-up (5 min) — Evaluation plan, monitoring, failure modes, and what you would do next.
45:00
Not started
Self-grade
Tick what you did. 0 of 8.
The model answer, step by step
Five steps, one tap each. The readout gives the step's answer and lists the numbers it uses; the same numbers are highlighted in the table underneath.
Customer-support agent, one tap per step
👉 Predict first, then tap. Every number below is computed from the assumptions in the table.
1. Framing
5 min
2. Architecture
10 min
3. Deep dive
15 min
4. Trade-offs
10 min
5. Wrap-up
5 min
Tap a step above.
Quantity
How it is computed
Value
Conversations per day
assumption
40,000
Share in the peak hour
assumption
12%
Model calls per conversation (tool loop)
assumption
4
Average input tokens per call
assumption
3,000
Average output tokens per call
assumption
250
Assumed price, $ per million input tokens
assumption
$3.00
Assumed price, $ per million output tokens
assumption
$15.00
Conversation length in seconds
assumption
360
Share resolved without a human
assumption
60%
Model cost per 1,000 conversations
calls × (in × pIn + out × pOut) ÷ 1e6 × 1000
$51.00
Model cost per 1,000 RESOLVED conversations
cost per 1,000 ÷ containment
$85.00
Peak conversation arrivals per second
conv × peak ÷ 3600
1.33
Concurrent live conversations at peak (Little's law)
arrival rate × duration
480
Input tokens per day
conv × calls × in
480,000,000
Takeaway. The headline figure — concurrent live conversations at peak (little's law) — is 480
(arrival rate × duration). Say the assumption, show the formula, then give the number.
The evaluation plan
Replay 300 real past conversations with tools mocked; grade resolution and action correctness.
Action audit: every write-tool call checked against policy in a test set of tricky requests (social engineering, prompt injection in customer text).
Online: containment, re-contact within 7 days (did it really resolve?), CSAT vs human baseline.
Regression: any prompt change must not lower the wrong-refund rate on the frozen set.