Document-processing pipeline

A Java batch job is sized by records per second. Here each 'record' is a page the model must read (prefill) and a JSON it must write (decode), and the two phases cost GPU time differently — so you size the fleet from GPU-seconds per page, not from a single throughput number.

45 min round 6 computed numbers 4-part eval plan 8-point rubric
⚠️ Planning numbers, not measurements. Hardware and model facts are dated on the numbers sheet; traffic, prices and efficiency are labelled assumptions. Replace them with your own measurements (the vLLM load-test exercise) before quoting them.

The prompt

Design a pipeline that turns 2 million scanned pages per day into validated JSON records.

Attempt it first: the 45-minute round

Set the timer, answer out loud or on paper, then score yourself against the rubric before you read the model answer below.

  1. Framing (5 min) — Ask questions, fix the scale and the latency/quality/cost targets, state assumptions with numbers.
  2. Architecture (10 min) — Draw the boxes end to end: data in, the model call, storage, serving, feedback.
  3. Deep dive (15 min) — Pick the hardest part and do the arithmetic: tokens/s, KV memory, QPS → replicas, cost per 1,000 requests.
  4. Trade-offs (10 min) — Name what you would trade (batch vs latency, quality vs cost, build vs buy) and what breaks.
  5. Wrap-up (5 min) — Evaluation plan, monitoring, failure modes, and what you would do next.
45:00
Not started

Self-grade

Tick what you did. 0 of 8.

The model answer, step by step

Five steps, one tap each. The readout gives the step's answer and lists the numbers it uses; the same numbers are highlighted in the table underneath.

Document-processing pipeline, one tap per step

👉 Predict first, then tap. Every number below is computed from the assumptions in the table.

1. Framing

5 min

2. Architecture

10 min

3. Deep dive

15 min

4. Trade-offs

10 min

5. Wrap-up

5 min

Tap a step above.
QuantityHow it is computedValue
Pages per dayassumption2,000,000
Processing window, secondsassumption28,800
Input tokens per page (image + prompt)assumption1,200
Output tokens per page (JSON)assumption500
Decode batchassumption64
Target GPU utilisationassumption80%
Share of pages reprocessed (validation fails)assumption3%
Pages per second to processpages ÷ window69.4
Prefill GPU-seconds per page (8B, 400 TFLOP/s achieved)2 × 8e9 × input ÷ 400e120.0480 s
Decode GPU-seconds per pageoutput × step ÷ batch0.0706 s
Total GPU-seconds per pageprefill + decode0.1186 s
GPUs incl. reprocessingpps × (1 + retry) × GPU-s per page ÷ utilisation11
GPU cost per 1,000 pages at $4/GPU-hourGPU-s per page × 1000 ÷ 3600 × $4$0.132
Takeaway. The headline figure — gpus incl. reprocessing — is 11 (pps × (1 + retry) × GPU-s per page ÷ utilisation). Say the assumption, show the formula, then give the number.

The evaluation plan

Go deeper on this site

Check yourself