Enterprise search / RAG at scale

A Java engineer reaches for a search index plus a service. RAG adds two twists: the index is vectors, so memory is the sizing question, and the answer must never cite a chunk the asker is not allowed to read — access control is part of retrieval, not a filter bolted on afterwards.

45 min round 9 computed numbers 4-part eval plan 8-point rubric
⚠️ Planning numbers, not measurements. Hardware and model facts are dated on the numbers sheet; traffic, prices and efficiency are labelled assumptions. Replace them with your own measurements (the vLLM load-test exercise) before quoting them.

The prompt

Design question-answering over a 50,000-person company's documents, with per-user access control.

Attempt it first: the 45-minute round

Set the timer, answer out loud or on paper, then score yourself against the rubric before you read the model answer below.

  1. Framing (5 min) — Ask questions, fix the scale and the latency/quality/cost targets, state assumptions with numbers.
  2. Architecture (10 min) — Draw the boxes end to end: data in, the model call, storage, serving, feedback.
  3. Deep dive (15 min) — Pick the hardest part and do the arithmetic: tokens/s, KV memory, QPS → replicas, cost per 1,000 requests.
  4. Trade-offs (10 min) — Name what you would trade (batch vs latency, quality vs cost, build vs buy) and what breaks.
  5. Wrap-up (5 min) — Evaluation plan, monitoring, failure modes, and what you would do next.
45:00
Not started

Self-grade

Tick what you did. 0 of 8.

The model answer, step by step

Five steps, one tap each. The readout gives the step's answer and lists the numbers it uses; the same numbers are highlighted in the table underneath.

Enterprise search / RAG at scale, one tap per step

👉 Predict first, then tap. Every number below is computed from the assumptions in the table.

1. Framing

5 min

2. Architecture

10 min

3. Deep dive

15 min

4. Trade-offs

10 min

5. Wrap-up

5 min

Tap a step above.
QuantityHow it is computedValue
Employeesassumption50,000
Questions per employee per dayassumption6
Share of daily questions in the peak hourassumption20%
Chunks in the corpusassumption20,000,000
Embedding dimensionassumption1,024
HNSW graph overhead factorassumption1.3
Prompt tokens (6 chunks × 400 + question + instructions)assumption2,700
Answer tokensassumption350
Batch per replicaassumption32
Peak questions per secondemp × qpd × peak ÷ 360016.7
Vector memory, fp16chunks × dim × 2 bytes41.0 GB
Vector memory, int8-quantisedchunks × dim × 1 byte20.5 GB
Int8 index in RAM with graphint8 × 1.326.6 GB
KV per request (8B, fp16)2 × 32 × 8 × 128 × seq × 2400 MB
Output tokens/s per 8B replica on one GPUbatch ÷ ((16 GB + batch × KV) ÷ 3.35 TB/s)3,723 tok/s
Output tokens/s needed at peakqps × answer tokens5,833 tok/s
Replicas incl. one spare⌈needed ÷ per-replica⌉ + 13
Generation cost per 1,000 questions at $4/GPU-hour$4 ÷ (tok/s × 3600 ÷ answer tokens) × 1000$0.104
Takeaway. The headline figure — int8 index in ram with graph — is 26.6 GB (int8 × 1.3). Say the assumption, show the formula, then give the number.

The evaluation plan

Go deeper on this site

Check yourself