In a Java IDE, completion is a local index lookup. A model completion is a network round trip that must feel local, so the design is dominated by latency: a small model near the user, context assembled fast, and prefix caching so you do not pay to re-read the same file on every keystroke.
45 min round7 computed numbers4-part eval plan8-point rubric
⚠️ Planning numbers, not measurements. Hardware and model facts
are dated on the numbers sheet; traffic, prices and efficiency are labelled
assumptions. Replace them with your own measurements (the vLLM load-test exercise) before quoting them.
The prompt
Design an IDE coding assistant with inline completions and a chat panel for 20,000 developers.
Attempt it first: the 45-minute round
Set the timer, answer out loud or on paper, then score yourself against the rubric before you read the model answer below.
Framing (5 min) — Ask questions, fix the scale and the latency/quality/cost targets, state assumptions with numbers.
Architecture (10 min) — Draw the boxes end to end: data in, the model call, storage, serving, feedback.
Deep dive (15 min) — Pick the hardest part and do the arithmetic: tokens/s, KV memory, QPS → replicas, cost per 1,000 requests.
Trade-offs (10 min) — Name what you would trade (batch vs latency, quality vs cost, build vs buy) and what breaks.
Wrap-up (5 min) — Evaluation plan, monitoring, failure modes, and what you would do next.
45:00
Not started
Self-grade
Tick what you did. 0 of 8.
The model answer, step by step
Five steps, one tap each. The readout gives the step's answer and lists the numbers it uses; the same numbers are highlighted in the table underneath.
Coding assistant, one tap per step
👉 Predict first, then tap. Every number below is computed from the assumptions in the table.
1. Framing
5 min
2. Architecture
10 min
3. Deep dive
15 min
4. Trade-offs
10 min
5. Wrap-up
5 min
Tap a step above.
Quantity
How it is computed
Value
Developers
assumption
20,000
Completion requests per developer per day
assumption
300
Share in the peak hour
assumption
12%
Context tokens per completion
assumption
1,500
Prefix-cache hit rate
assumption
60%
Completion length, tokens
assumption
40
Decode batch
assumption
32
Target utilisation
assumption
60%
Peak completions per second
devs × rpd × peak ÷ 3600
200
Prefill GPU-seconds (7B ≈ 7e9 params, cache hit skips)
(1 − hit) × 2 × 7e9 × ctx ÷ 400e12
0.0210 s
Decode GPU-seconds per completion
tokens × step ÷ batch (14 GB weights)
0.0076 s
GPU-seconds per completion
prefill + decode
0.0286 s
GPUs
qps × GPU-s ÷ utilisation
10
GPU cost per 1,000 completions at $4/GPU-hour
GPU-s × 1000 ÷ 3600 × $4
$0.032
GPUs with no prefix caching
same, hit = 0
21
Takeaway. The headline figure — gpus — is 10
(qps × GPU-s ÷ utilisation). Say the assumption, show the formula, then give the number.
The evaluation plan
Offline: fill-in-the-middle exact/compile-pass rate on held-out repos, refreshed so it is not memorised.
Functional: unit tests pass after applying suggested edits on a bench of 200 tasks.
Online: acceptance rate, kept-after-30-seconds rate and latency percentiles per region with a control group.
Safety: secret and licence-leak scans on a sample of outputs.