On-device assistant, privacy-first

A Java server assumes you can add memory. A phone has a fixed few gigabytes shared with the OS, and a battery. The design is constrained from the bottom: bytes of weights and KV cache, memory bandwidth, and joules per token decide what can run at all.

45 min round 7 computed numbers 4-part eval plan 8-point rubric
⚠️ Planning numbers, not measurements. Hardware and model facts are dated on the numbers sheet; traffic, prices and efficiency are labelled assumptions. Replace them with your own measurements (the vLLM load-test exercise) before quoting them.

The prompt

Design an assistant that runs on a phone and keeps personal data on the device.

Attempt it first: the 45-minute round

Set the timer, answer out loud or on paper, then score yourself against the rubric before you read the model answer below.

  1. Framing (5 min) — Ask questions, fix the scale and the latency/quality/cost targets, state assumptions with numbers.
  2. Architecture (10 min) — Draw the boxes end to end: data in, the model call, storage, serving, feedback.
  3. Deep dive (15 min) — Pick the hardest part and do the arithmetic: tokens/s, KV memory, QPS → replicas, cost per 1,000 requests.
  4. Trade-offs (10 min) — Name what you would trade (batch vs latency, quality vs cost, build vs buy) and what breaks.
  5. Wrap-up (5 min) — Evaluation plan, monitoring, failure modes, and what you would do next.
45:00
Not started

Self-grade

Tick what you did. 0 of 8.

The model answer, step by step

Five steps, one tap each. The readout gives the step's answer and lists the numbers it uses; the same numbers are highlighted in the table underneath.

On-device assistant, privacy-first, one tap per step

👉 Predict first, then tap. Every number below is computed from the assumptions in the table.

1. Framing

5 min

2. Architecture

10 min

3. Deep dive

15 min

4. Trade-offs

10 min

5. Wrap-up

5 min

Tap a step above.
QuantityHow it is computedValue
Model parameters (assumed Llama-3.2-3B-like: 28 layers, 8 KV heads, head dim 128)assumption3,000,000,000
Bytes per weight (int4)assumption0.5
Context tokensassumption4,096
Bytes per KV value (int8)assumption1
Phone memory bandwidth, GB/s (assumed)assumption50
Battery, Wh (assumed)assumption15
Sustained draw while generating, W (assumed)assumption5
Usersassumption5,000,000
Requests per user per dayassumption10
Share falling back to the cloudassumption15%
Cloud cost per 1,000 fallback requests (assumed)assumption$2.00
Weight memoryparams × 0.5 bytes1.50 GB
KV cache at full context2 × 28 × 8 × 128 × seq × 1 byte235 MB
Model + KV memoryweights + KV1.73 GB
Decode speed (memory-bound)bandwidth ÷ weight bytes33 tok/s
Tokens generated per 1% of battery(Wh × 3600 × 1% ÷ W) × tok/s3,600
Cloud fallback requests per dayusers × rpd × fallback7,500,000
Cloud fallback cost per dayrequests ÷ 1000 × $/1k$15,000
Takeaway. The headline figure — model + kv memory — is 1.73 GB (weights + KV). Say the assumption, show the formula, then give the number.

The evaluation plan

Go deeper on this site

Check yourself