A Java team gates a release on unit tests that are deterministic. LLM outputs are not, so an eval run is a statistics problem: how many cases before a three-point drop is distinguishable from noise, and who judges the answers — a rubric, a model, or a human.
45 min round6 computed numbers4-part eval plan8-point rubric
⚠️ Planning numbers, not measurements. Hardware and model facts
are dated on the numbers sheet; traffic, prices and efficiency are labelled
assumptions. Replace them with your own measurements (the vLLM load-test exercise) before quoting them.
The prompt
Design an evaluation platform so 30 teams can gate prompt and model changes on quality.
Attempt it first: the 45-minute round
Set the timer, answer out loud or on paper, then score yourself against the rubric before you read the model answer below.
Framing (5 min) — Ask questions, fix the scale and the latency/quality/cost targets, state assumptions with numbers.
Architecture (10 min) — Draw the boxes end to end: data in, the model call, storage, serving, feedback.
Deep dive (15 min) — Pick the hardest part and do the arithmetic: tokens/s, KV memory, QPS → replicas, cost per 1,000 requests.
Trade-offs (10 min) — Name what you would trade (batch vs latency, quality vs cost, build vs buy) and what breaks.
Wrap-up (5 min) — Evaluation plan, monitoring, failure modes, and what you would do next.
45:00
Not started
Self-grade
Tick what you did. 0 of 8.
The model answer, step by step
Five steps, one tap each. The readout gives the step's answer and lists the numbers it uses; the same numbers are highlighted in the table underneath.
Eval platform, one tap per step
👉 Predict first, then tap. Every number below is computed from the assumptions in the table.
1. Framing
5 min
2. Architecture
10 min
3. Deep dive
15 min
4. Trade-offs
10 min
5. Wrap-up
5 min
Tap a step above.
Quantity
How it is computed
Value
Current pass rate
assumption
85%
Regression to detect
assumption
82%
z for one-sided α = 0.05
assumption
1.645
z for 80% power
assumption
0.842
Candidate call: input tokens
assumption
1,500
Candidate call: output tokens
assumption
400
Judge call: input tokens
assumption
2,000
Judge call: output tokens
assumption
150
Assumed price, $ per million input tokens
assumption
$3.00
Assumed price, $ per million output tokens
assumption
$15.00
Eval runs per day across teams
assumption
60
Concurrent calls
assumption
50
Seconds per candidate+judge pair
assumption
6
Cases needed
(zα+zβ)² × (p1q1 + p2q2) ÷ (p1−p2)²
1,891
Cost of one case (candidate + judge)
(in × pIn + out × pOut) ÷ 1e6, both calls
$0.019
Cost of one run at that size
cases × cost per case
$35.46
Cost per day
runs × run cost
$2,127
Wall-clock per run
cases × seconds per pair ÷ concurrency
227 s
Cost per 1,000 cases
1000 × cost per case
$18.75
Takeaway. The headline figure — cases needed — is 1,891
((zα+zβ)² × (p1q1 + p2q2) ÷ (p1−p2)²). Say the assumption, show the formula, then give the number.
The evaluation plan
Meta-eval: judge vs 200 human-labelled answers, report agreement (Cohen's kappa) and re-run monthly.
Seed known regressions (a deliberately worse prompt) and confirm the gate catches them; seed a no-op change and confirm it does not.
Track flake rate: run the same config twice and compare scores.
Coverage: each team's set must include their known past incidents.