You build two pieces of pure math and one small async tool: the standard
transformer KV-cache memory formula, a nearest-rank latency percentile, and an
asyncio + httpx load generator that hits any
OpenAI-compatible /v1/completions-shaped endpoint at increasing concurrency
and reports p50/p95 per level — the same two questions packaging
& serving and the SageMaker endpoint left open:
how much GPU memory does this batch and context length cost, and how does latency behave
as load rises. It runs on your Mac (and, for the real deliverable, against a local vLLM
server) rather than in the browser because Pyodide has no real sockets and no GPU.
kv_cache_bytes(layers, heads_kv, head_dim, seq, batch, dtype) —
2 * layers * heads_kv * head_dim * seq * batch * dtype_bytes. Note
heads_kv: a GQA model's KV-head count, not its query-head count.percentile(samples, p) — nearest-rank percentile; must raise on an empty
sequence or a p outside [0, 100].run_at_concurrency(...) / run_load_test(...) — fire requests
at an endpoint through an asyncio.Semaphore-bounded pool and report p50/p95
per concurrency level; format_report(...) renders that as markdown.The tests are ordinary pytest and ship in the public folder with the starter — read them first; the names below are the check list. Solutions are not published.
Needs git. uv installs the right Python itself, so nothing else is required.
# once, anywhere on your machine
git clone https://github.com/theDocWho/ai-ml-roadmap.git
cd ai-ml-roadmap
No git? Download the ZIP, unzip it, and cd into the unzipped folder instead.
From the repo root:
# one-time: uv (https://docs.astral.sh/uv/) manages the venv and pins Python ≥ 3.12 cd exercises/ex-vllm-load && uv sync && uv run pytest -q # the real deliverable: run it for real against a local OpenAI-compatible server # (vLLM, llama.cpp's server, Ollama's /v1 route — anything that speaks the same JSON) uv run python -c " import asyncio from vllm_load import format_report, kv_cache_bytes, run_load_test report = asyncio.run(run_load_test( 'http://localhost:8000', concurrency_levels=[1, 2, 4, 8], requests_per_level=20, )) print(format_report(report)) print(kv_cache_bytes(layers=32, heads_kv=8, head_dim=128, seq=8192, batch=1, dtype='fp16')) " # the same bar the reference solution clears uv run ruff check . && uv run mypy src
Done when uv run pytest -q prints 5 passed. That
suite certifies the KV-cache arithmetic, the percentile function, the report's shape, and
that the load generator's concurrency plumbing produces more contention at higher
concurrency against a fake, in-process server (tests/conftest.py) — it does
not run against a real GPU. The actual deliverable for the checkpoint is the
format_report(...) table and the kv_cache_bytes(...) figure from a
real run against your own vLLM server, pasted in — the numbers below (the RTX 3060 run) are
what that looks like.
test_kv_cache_bytes_llama3_8b_8k_batch1_fp16 — Llama-3-8B's published
architecture (32 layers, 8 KV heads, head_dim 128) at 8k context, fp16, batch 1 must come
out to exactly 1,073,741,824 bytes (1 GiB) — and the same cache at int8 to half
that, 536,870,912 bytes, so dtype has to be looked up, not assumed.test_kv_cache_bytes_llama3_8b_8k_batch16_fp16 — the same config at batch
16 must come out to exactly 17,179,869,184 bytes (16 GiB) — 16× the batch-1 figure.test_percentile_matches_hand_computed_p50_and_p95 — against the
values 10, 20, ..., 100 handed over unsorted (latencies arrive in
completion order): p50 = 50, p95 = 100.test_format_report_has_header_and_one_row_per_level — the markdown table
has a header row naming concurrency/p50/p95 and exactly one data row per concurrency
level, in order.test_p50_rises_monotonically_with_concurrency — a real run of
run_load_test against a fake OpenAI-compatible server that starts inside the
test (its simulated latency grows with how many requests it is serving at once) must show
p50 rising as concurrency rises through [1, 4, 16].percentileThis is self-attestation — the site cannot see your terminal, so the box and the button are you telling The Path the suite went green on your machine.
kv_cache_bytes — write the six-factor product one line at a time
before collapsing it into one expression; a missing factor of 2 (K and V are both
cached) or swapping the query-head count for heads_kv are the two classic
mistakes.percentile — sort first. The rank is
ceil(p / 100 * n), 1-indexed, so the list index is
rank - 1; dropping either the ceiling or the - 1 shifts every
answer by one sample.run_at_concurrency — hold an asyncio.Semaphore(concurrency)
across each individual request (inside the coroutine that awaits it), not just around the
call to asyncio.gather; launch all n_requests coroutines through
one gather and let the semaphore throttle them.uv run pytest -q -x --tb=short stops at the first
failure and shows the assertion that tripped; the message names the behaviour, not the fix.