A Java batch job is sized by records per second. Here each 'record' is a page the model must read (prefill) and a JSON it must write (decode), and the two phases cost GPU time differently — so you size the fleet from GPU-seconds per page, not from a single throughput number.
45 min round6 computed numbers4-part eval plan8-point rubric
⚠️ Planning numbers, not measurements. Hardware and model facts
are dated on the numbers sheet; traffic, prices and efficiency are labelled
assumptions. Replace them with your own measurements (the vLLM load-test exercise) before quoting them.
The prompt
Design a pipeline that turns 2 million scanned pages per day into validated JSON records.
Attempt it first: the 45-minute round
Set the timer, answer out loud or on paper, then score yourself against the rubric before you read the model answer below.
Framing (5 min) — Ask questions, fix the scale and the latency/quality/cost targets, state assumptions with numbers.
Architecture (10 min) — Draw the boxes end to end: data in, the model call, storage, serving, feedback.
Deep dive (15 min) — Pick the hardest part and do the arithmetic: tokens/s, KV memory, QPS → replicas, cost per 1,000 requests.
Trade-offs (10 min) — Name what you would trade (batch vs latency, quality vs cost, build vs buy) and what breaks.
Wrap-up (5 min) — Evaluation plan, monitoring, failure modes, and what you would do next.
45:00
Not started
Self-grade
Tick what you did. 0 of 8.
The model answer, step by step
Five steps, one tap each. The readout gives the step's answer and lists the numbers it uses; the same numbers are highlighted in the table underneath.
Document-processing pipeline, one tap per step
👉 Predict first, then tap. Every number below is computed from the assumptions in the table.
1. Framing
5 min
2. Architecture
10 min
3. Deep dive
15 min
4. Trade-offs
10 min
5. Wrap-up
5 min
Tap a step above.
Quantity
How it is computed
Value
Pages per day
assumption
2,000,000
Processing window, seconds
assumption
28,800
Input tokens per page (image + prompt)
assumption
1,200
Output tokens per page (JSON)
assumption
500
Decode batch
assumption
64
Target GPU utilisation
assumption
80%
Share of pages reprocessed (validation fails)
assumption
3%
Pages per second to process
pages ÷ window
69.4
Prefill GPU-seconds per page (8B, 400 TFLOP/s achieved)
2 × 8e9 × input ÷ 400e12
0.0480 s
Decode GPU-seconds per page
output × step ÷ batch
0.0706 s
Total GPU-seconds per page
prefill + decode
0.1186 s
GPUs incl. reprocessing
pps × (1 + retry) × GPU-s per page ÷ utilisation
11
GPU cost per 1,000 pages at $4/GPU-hour
GPU-s per page × 1000 ÷ 3600 × $4
$0.132
Takeaway. The headline figure — gpus incl. reprocessing — is 11
(pps × (1 + retry) × GPU-s per page ÷ utilisation). Say the assumption, show the formula, then give the number.
The evaluation plan
Field-level exact match on 1,000 human-labelled pages per document type (invoice, receipt, form).
Business-rule checks: line items sum to the total; dates parse; IDs match a registry.
Shadow-run any model change on last week's pages and diff the JSON before switching.
Human-review sample of 1% of accepted pages each day to catch confident errors.