A Java engineer's search is an inverted index over words. Multimodal search embeds text and images into one vector space and does nearest-neighbour lookup; the sizing question moves from disk to RAM, and compression trades memory for recall.
45 min round8 computed numbers4-part eval plan8-point rubric
⚠️ Planning numbers, not measurements. Hardware and model facts
are dated on the numbers sheet; traffic, prices and efficiency are labelled
assumptions. Replace them with your own measurements (the vLLM load-test exercise) before quoting them.
The prompt
Design search where users type text and find images among 100 million photos.
Attempt it first: the 45-minute round
Set the timer, answer out loud or on paper, then score yourself against the rubric before you read the model answer below.
Framing (5 min) — Ask questions, fix the scale and the latency/quality/cost targets, state assumptions with numbers.
Architecture (10 min) — Draw the boxes end to end: data in, the model call, storage, serving, feedback.
Deep dive (15 min) — Pick the hardest part and do the arithmetic: tokens/s, KV memory, QPS → replicas, cost per 1,000 requests.
Trade-offs (10 min) — Name what you would trade (batch vs latency, quality vs cost, build vs buy) and what breaks.
Wrap-up (5 min) — Evaluation plan, monitoring, failure modes, and what you would do next.
45:00
Not started
Self-grade
Tick what you did. 0 of 8.
The model answer, step by step
Five steps, one tap each. The readout gives the step's answer and lists the numbers it uses; the same numbers are highlighted in the table underneath.
Multimodal search, one tap per step
👉 Predict first, then tap. Every number below is computed from the assumptions in the table.
1. Framing
5 min
2. Architecture
10 min
3. Deep dive
15 min
4. Trade-offs
10 min
5. Wrap-up
5 min
Tap a step above.
Quantity
How it is computed
Value
Images
assumption
100,000,000
Embedding dimension
assumption
768
Product-quantised bytes per vector
assumption
64
Index overhead factor
assumption
1.3
Peak queries per second
assumption
1,500
Queries per second one index replica serves (assumed)
assumption
300
Usable RAM per node, GB (assumed)
assumption
48
Images embedded per GPU-second (assumed)
assumption
400
Raw vectors, fp16
images × dim × 2
153.6 GB
Raw vectors, int8
images × dim × 1
76.8 GB
Product-quantised vectors
images × 64 bytes
6.4 GB
PQ index in RAM with overhead
PQ × 1.3
8.3 GB
Shards by memory (fp16 raw)
⌈raw ÷ RAM per node⌉
4
Replicas to serve peak QPS
⌈qps ÷ per-replica⌉
5
One-off embedding GPU-hours
images ÷ rate ÷ 3600
69.4
One-off embedding cost at $4/GPU-hour
GPU-hours × $4
$278
Takeaway. The headline figure — pq index in ram with overhead — is 8.3 GB
(PQ × 1.3). Say the assumption, show the formula, then give the number.
The evaluation plan
Recall of the ANN index versus exact search on 10,000 sampled queries (target ≥ 95% at k = 10).
Relevance: 500 labelled query→image pairs, nDCG@10, plus a slice per language and per rare concept.
Bias and safety: audit results for sensitive queries; block unsafe content at index time.
Online: click-through on the top 3 and reformulation rate.