# ex-vllm-load — KV-cache sizing, a latency percentile, and a real load run against vLLM

Serving an LLM well means answering two questions before you ever send a request: how much
GPU memory will the KV cache for this batch and context length cost, and how does latency
behave as concurrency rises. This exercise builds both: a pure formula for the first, and a
small `asyncio` + `httpx` load generator for the second, against any OpenAI-compatible
`/v1/completions`-shaped endpoint (vLLM, llama.cpp's server, Ollama's OpenAI-compatible
route — anything that speaks the same JSON).

## What to implement

`src/vllm_load/kv_cache.py`, `src/vllm_load/stats.py`, `src/vllm_load/loadtest.py` — the
signatures and docstrings are the contract; fill in the bodies (all currently `...`):

- `kv_cache_bytes(layers, heads_kv, head_dim, seq, batch, dtype)` — the standard transformer
  KV-cache formula: `2 (K and V) * layers * heads_kv * head_dim * seq * batch * dtype_bytes`.
  `heads_kv` is the GQA KV-head count, not the query-head count — the two differ on any
  modern model (Llama 3's 8B has 32 query heads but only 8 KV heads).
- `percentile(samples, p)` — nearest-rank percentile (module docstring has the exact rule).
- `run_at_concurrency(...)` / `run_load_test(...)` — fire requests at an endpoint, bounded by
  an `asyncio.Semaphore`, and report p50/p95 per concurrency level.
- `format_report(report)` — that report as a markdown table.

## Run it

```
cd exercises/ex-vllm-load && uv sync && uv run pytest -q
```

(`python3 -m venv .venv && .venv/bin/pip install -q httpx pytest pytest-asyncio && .venv/bin/python -m pytest -q`
works too if you don't have `uv`.)

## The checks (5)

- `test_kv_cache_bytes_llama3_8b_8k_batch1_fp16` / `..._batch16_fp16` — the formula against
  Llama-3-8B's published architecture (32 layers, 8 KV heads, head_dim 128) at 8k context;
  the arithmetic is in the test's comments. The batch-1 test also asks for `int8` (half of fp16).
- `test_percentile_matches_hand_computed_p50_and_p95` — against the hand-computed values
  `10, 20, ..., 100`, passed in unsorted.
- `test_format_report_has_header_and_one_row_per_level` — the markdown table's shape.
- `test_p50_rises_monotonically_with_concurrency` — a real run of `run_load_test` against a
  fake OpenAI-compatible server (`tests/conftest.py`, starts fresh in the test) whose
  simulated latency grows with how many requests it is serving at once, so p50 must rise as
  concurrency does.

**What the test suite certifies, and what it does not:** these five checks certify the
formula's arithmetic, the percentile function, the report's shape, and that the load
generator's concurrency plumbing actually produces more contention at higher concurrency
against a fake, in-process server. None of that is the real deliverable. The real deliverable
is running `run_load_test` for real against a vLLM (or other OpenAI-compatible) server —
ideally on the RTX 3060 the flagship targets — at a few concurrency levels, and pasting the
resulting `format_report(...)` table plus the `kv_cache_bytes(...)` figure for your model and
context length into the day's checkpoint. A green `pytest -q` here proves the tool works; it
does not replace the real run.

```python
import asyncio
from vllm_load import format_report, kv_cache_bytes, run_load_test

report = asyncio.run(run_load_test(
    "http://localhost:8000",
    concurrency_levels=[1, 2, 4, 8],
    requests_per_level=20,
    path="/v1/completions",
    payload={"model": "your-model", "prompt": "Explain KV caching in one paragraph.", "max_tokens": 64},
))
print(format_report(report))
print(kv_cache_bytes(layers=32, heads_kv=8, head_dim=128, seq=8192, batch=1, dtype="fp16"))
```

## If you get stuck

- **`kv_cache_bytes`** — write out the six-factor product one line at a time before
  collapsing it; a missing factor of 2 (K and V are both cached) or swapping `heads_kv` for
  the query-head count are the two classic mistakes, and both are in `traps/` for a reason.
- **`percentile`** — sort first. The rank formula is `ceil(p / 100 * n)`, **1-indexed**, so
  the list index is `rank - 1`; forgetting either the ceiling or the `- 1` shifts every
  answer by one sample.
- **`run_at_concurrency`** — bound concurrency with `asyncio.Semaphore(concurrency)` held
  across each request, not just around the `gather` call; launch all `n_requests` coroutines
  through `asyncio.gather`, and let the semaphore do the throttling.
- **`format_report`** — the test only checks for `"| {level} |"` verbatim in each row and
  the three header words; any markdown table with a header, a separator and one row per
  level in order satisfies it.
- **Reading a red row** — `uv run pytest -q -x --tb=short` stops at the first failure and
  shows the assertion that tripped; the message names the behaviour, not the fix.

## Files

- [`exercises/ex-vllm-load/README.md`](README.md) — this brief, offline
- [`pyproject.toml`](pyproject.toml) — deps, ruff and mypy config
- [`src/vllm_load/kv_cache.py`](src/vllm_load/kv_cache.py) — `kv_cache_bytes`, the starter you edit
- [`src/vllm_load/stats.py`](src/vllm_load/stats.py) — `percentile`, the starter you edit
- [`src/vllm_load/loadtest.py`](src/vllm_load/loadtest.py) — the load generator, the starter you edit
- [`tests/conftest.py`](tests/conftest.py) — the fake OpenAI-compatible server fixture
- [`tests/test_kv_cache.py`](tests/test_kv_cache.py), [`tests/test_percentile.py`](tests/test_percentile.py), [`tests/test_report.py`](tests/test_report.py), [`tests/test_load_run.py`](tests/test_load_run.py) — the checks

Done when `uv run pytest -q` prints `5 passed`, then press "mark done" on the page — after
you've also run it for real against a live server and pasted the numbers into the checkpoint.
