Exercise ex-vllm-load — size the KV cache, then load-test the endpoint for real

You build two pieces of pure math and one small async tool: the standard transformer KV-cache memory formula, a nearest-rank latency percentile, and an asyncio + httpx load generator that hits any OpenAI-compatible /v1/completions-shaped endpoint at increasing concurrency and reports p50/p95 per level — the same two questions packaging & serving and the SageMaker endpoint left open: how much GPU memory does this batch and context length cost, and how does latency behave as load rises. It runs on your Mac (and, for the real deliverable, against a local vLLM server) rather than in the browser because Pyodide has no real sockets and no GPU.

~120 minruns locally · uv + pytest 5 checksex-vllm-load

What you're building

The tests are ordinary pytest and ship in the public folder with the starter — read them first; the names below are the check list. Solutions are not published.

Get the repo (once)

Needs git. uv installs the right Python itself, so nothing else is required.

# once, anywhere on your machine
git clone https://github.com/theDocWho/ai-ml-roadmap.git
cd ai-ml-roadmap

No git? Download the ZIP, unzip it, and cd into the unzipped folder instead.

Run it

From the repo root:

# one-time: uv (https://docs.astral.sh/uv/) manages the venv and pins Python ≥ 3.12
cd exercises/ex-vllm-load && uv sync && uv run pytest -q

# the real deliverable: run it for real against a local OpenAI-compatible server
# (vLLM, llama.cpp's server, Ollama's /v1 route — anything that speaks the same JSON)
uv run python -c "
import asyncio
from vllm_load import format_report, kv_cache_bytes, run_load_test

report = asyncio.run(run_load_test(
    'http://localhost:8000',
    concurrency_levels=[1, 2, 4, 8],
    requests_per_level=20,
))
print(format_report(report))
print(kv_cache_bytes(layers=32, heads_kv=8, head_dim=128, seq=8192, batch=1, dtype='fp16'))
"

# the same bar the reference solution clears
uv run ruff check . && uv run mypy src

Done when uv run pytest -q prints 5 passed. That suite certifies the KV-cache arithmetic, the percentile function, the report's shape, and that the load generator's concurrency plumbing produces more contention at higher concurrency against a fake, in-process server (tests/conftest.py) — it does not run against a real GPU. The actual deliverable for the checkpoint is the format_report(...) table and the kv_cache_bytes(...) figure from a real run against your own vLLM server, pasted in — the numbers below (the RTX 3060 run) are what that looks like.

The checks

Files

This is self-attestation — the site cannot see your terminal, so the box and the button are you telling The Path the suite went green on your machine.

If you get stuck