A Java server assumes you can add memory. A phone has a fixed few gigabytes shared with the OS, and a battery. The design is constrained from the bottom: bytes of weights and KV cache, memory bandwidth, and joules per token decide what can run at all.
45 min round7 computed numbers4-part eval plan8-point rubric
⚠️ Planning numbers, not measurements. Hardware and model facts
are dated on the numbers sheet; traffic, prices and efficiency are labelled
assumptions. Replace them with your own measurements (the vLLM load-test exercise) before quoting them.
The prompt
Design an assistant that runs on a phone and keeps personal data on the device.
Attempt it first: the 45-minute round
Set the timer, answer out loud or on paper, then score yourself against the rubric before you read the model answer below.
Framing (5 min) — Ask questions, fix the scale and the latency/quality/cost targets, state assumptions with numbers.
Architecture (10 min) — Draw the boxes end to end: data in, the model call, storage, serving, feedback.
Deep dive (15 min) — Pick the hardest part and do the arithmetic: tokens/s, KV memory, QPS → replicas, cost per 1,000 requests.
Trade-offs (10 min) — Name what you would trade (batch vs latency, quality vs cost, build vs buy) and what breaks.
Wrap-up (5 min) — Evaluation plan, monitoring, failure modes, and what you would do next.
45:00
Not started
Self-grade
Tick what you did. 0 of 8.
The model answer, step by step
Five steps, one tap each. The readout gives the step's answer and lists the numbers it uses; the same numbers are highlighted in the table underneath.
On-device assistant, privacy-first, one tap per step
👉 Predict first, then tap. Every number below is computed from the assumptions in the table.
1. Framing
5 min
2. Architecture
10 min
3. Deep dive
15 min
4. Trade-offs
10 min
5. Wrap-up
5 min
Tap a step above.
Quantity
How it is computed
Value
Model parameters (assumed Llama-3.2-3B-like: 28 layers, 8 KV heads, head dim 128)
assumption
3,000,000,000
Bytes per weight (int4)
assumption
0.5
Context tokens
assumption
4,096
Bytes per KV value (int8)
assumption
1
Phone memory bandwidth, GB/s (assumed)
assumption
50
Battery, Wh (assumed)
assumption
15
Sustained draw while generating, W (assumed)
assumption
5
Users
assumption
5,000,000
Requests per user per day
assumption
10
Share falling back to the cloud
assumption
15%
Cloud cost per 1,000 fallback requests (assumed)
assumption
$2.00
Weight memory
params × 0.5 bytes
1.50 GB
KV cache at full context
2 × 28 × 8 × 128 × seq × 1 byte
235 MB
Model + KV memory
weights + KV
1.73 GB
Decode speed (memory-bound)
bandwidth ÷ weight bytes
33 tok/s
Tokens generated per 1% of battery
(Wh × 3600 × 1% ÷ W) × tok/s
3,600
Cloud fallback requests per day
users × rpd × fallback
7,500,000
Cloud fallback cost per day
requests ÷ 1000 × $/1k
$15,000
Takeaway. The headline figure — model + kv memory — is 1.73 GB
(weights + KV). Say the assumption, show the formula, then give the number.
The evaluation plan
Quality: the same task suite run on-device (quantised) and in fp16 to measure the quantisation loss per task.
Device matrix: latency, memory and thermal tests on the lowest-spec supported phones.
Privacy: a network-traffic audit proving no personal data leaves in local mode.
Routing: precision/recall of the local-vs-cloud router against human judgment.