A Java system would filter with rules and a database of banned terms. LLM moderation is powerful but far too expensive to run on everything, so the design is a cascade — and base rates matter: when only 1% of posts violate policy, even a good classifier produces mostly false alarms.
45 min round6 computed numbers4-part eval plan8-point rubric
⚠️ Planning numbers, not measurements. Hardware and model facts
are dated on the numbers sheet; traffic, prices and efficiency are labelled
assumptions. Replace them with your own measurements (the vLLM load-test exercise) before quoting them.
The prompt
Design moderation for 50 million posts a day.
Attempt it first: the 45-minute round
Set the timer, answer out loud or on paper, then score yourself against the rubric before you read the model answer below.
Framing (5 min) — Ask questions, fix the scale and the latency/quality/cost targets, state assumptions with numbers.
Architecture (10 min) — Draw the boxes end to end: data in, the model call, storage, serving, feedback.
Deep dive (15 min) — Pick the hardest part and do the arithmetic: tokens/s, KV memory, QPS → replicas, cost per 1,000 requests.
Trade-offs (10 min) — Name what you would trade (batch vs latency, quality vs cost, build vs buy) and what breaks.
Wrap-up (5 min) — Evaluation plan, monitoring, failure modes, and what you would do next.
45:00
Not started
Self-grade
Tick what you did. 0 of 8.
The model answer, step by step
Five steps, one tap each. The readout gives the step's answer and lists the numbers it uses; the same numbers are highlighted in the table underneath.
Content moderation, one tap per step
👉 Predict first, then tap. Every number below is computed from the assumptions in the table.
1. Framing
5 min
2. Architecture
10 min
3. Deep dive
15 min
4. Trade-offs
10 min
5. Wrap-up
5 min
Tap a step above.
Quantity
How it is computed
Value
Posts per day
assumption
50,000,000
Peak-to-mean traffic ratio
assumption
3
Share decided by the cheap classifier
assumption
92%
LLM prompt tokens (post + policy)
assumption
600
LLM output tokens (label + reason)
assumption
20
Decode batch
assumption
32
Utilisation
assumption
60%
Share sent to human review
assumption
0.5%
Items one reviewer handles per day
assumption
400
Base rate of violating posts
assumption
1%
Classifier recall
assumption
95%
Classifier false-positive rate
assumption
2%
Peak posts per second
posts ÷ 86400 × peak ratio
1,737
Peak posts reaching the LLM
qps × (1 − decided)
139
GPU-seconds per LLM call (8B)
prefill + decode
0.0275 s
GPUs
LLM qps × GPU-s ÷ utilisation
7
Human reviewers needed
posts × human share ÷ per reviewer
625
Precision of the cheap classifier's flags
TP ÷ (TP + FP) at 1% base rate
32.4%
Takeaway. The headline figure — human reviewers needed — is 625
(posts × human share ÷ per reviewer). Say the assumption, show the formula, then give the number.
The evaluation plan
Audited random sample per policy (labelled by trained reviewers) to estimate real precision and recall, not just those on a curated set.
Adversarial set: obfuscated slurs, images with text, multilingual and coded language, refreshed monthly.
Fairness: error rates sliced by language and dialect.