In a Java recommender, ranking is a cheap model scoring thousands of candidates. An LLM is far more expensive per item, so it can only touch a short list — the design is a funnel, and the number to compute is how much of the traffic can afford the expensive stage.
45 min round8 computed numbers4-part eval plan8-point rubric
⚠️ Planning numbers, not measurements. Hardware and model facts
are dated on the numbers sheet; traffic, prices and efficiency are labelled
assumptions. Replace them with your own measurements (the vLLM load-test exercise) before quoting them.
The prompt
Add an LLM reranker to a feed for 20 million daily users without breaking the 150 ms latency budget.
Attempt it first: the 45-minute round
Set the timer, answer out loud or on paper, then score yourself against the rubric before you read the model answer below.
Framing (5 min) — Ask questions, fix the scale and the latency/quality/cost targets, state assumptions with numbers.
Architecture (10 min) — Draw the boxes end to end: data in, the model call, storage, serving, feedback.
Deep dive (15 min) — Pick the hardest part and do the arithmetic: tokens/s, KV memory, QPS → replicas, cost per 1,000 requests.
Trade-offs (10 min) — Name what you would trade (batch vs latency, quality vs cost, build vs buy) and what breaks.
Wrap-up (5 min) — Evaluation plan, monitoring, failure modes, and what you would do next.
45:00
Not started
Self-grade
Tick what you did. 0 of 8.
The model answer, step by step
Five steps, one tap each. The readout gives the step's answer and lists the numbers it uses; the same numbers are highlighted in the table underneath.
LLM-powered recommendations, one tap per step
👉 Predict first, then tap. Every number below is computed from the assumptions in the table.
1. Framing
5 min
2. Architecture
10 min
3. Deep dive
15 min
4. Trade-offs
10 min
5. Wrap-up
5 min
Tap a step above.
Quantity
How it is computed
Value
Daily active users
assumption
20,000,000
Sessions per user per day
assumption
6
Share of sessions sent to the LLM stage
assumption
25%
Share in the peak hour
assumption
10%
Prompt tokens (20 items × 40 + user summary)
assumption
1,000
Output tokens (one relevance-score token; the 20 items are scored as parallel sequences)
assumption
1
Decode batch
assumption
32
Target utilisation
assumption
60%
Peak LLM reranks per second
dau × sessions × share × peak ÷ 3600
833
Prefill GPU-seconds (8B)
2 × 8e9 × input ÷ 400e12
0.0400 s
Decode GPU-seconds
output × step ÷ batch
0.0002 s
GPUs
qps × (prefill + decode) ÷ utilisation
56
Cost per 1,000 reranks at $4/GPU-hour
(prefill + decode) × 1000 ÷ 3600 × $4
$0.045
LLM-stage latency of one rerank
prefill + one score step
46 ms
Latency if it generated a 30-token ranked list instead
prefill + 30 × step
222 ms
GPUs if every session used it
GPUs ÷ share
224
Takeaway. The headline figure — gpus — is 56
(qps × (prefill + decode) ÷ utilisation). Say the assumption, show the formula, then give the number.
The evaluation plan
Offline: rank-metrics (nDCG) of LLM vs cheap ranker on logged sessions, keeping time order to avoid leakage.
Online: A/B with a holdback for watch time or conversion, and separate latency guardrails.
Slice by new users, long-tail items and languages to find where it hurts.
Safety: audit the top-ranked items for policy-violating content.