# py-07-async-token-bucket — an async LLM client that respects both limits

Every RAG or agent service has exactly this component: a client that calls a model API under
two limits the provider imposes at once — requests per second *and* connections in flight —
retries a transient failure once, and gives up cleanly when its caller cancels. You build it
against a mock transport here; flagship F1 (the RAG service) swaps in a real one. The brief is
the same text as the page (`illustrated/0-python/exercise-async-token-bucket.html`).

This one runs on your Mac, not in the browser: it needs a real asyncio event loop with tasks,
locks and cancellation, which Pyodide's single-threaded loop does not exercise faithfully.

## What to implement

`src/llmclient/bucket.py` — `TokenBucket(rate, capacity, *, clock, sleep)`:

- starts full; refills `rate` tokens per second **continuously, capped at `capacity`** (an hour
  idle must not earn an hour of burst);
- `async acquire(n=1)` waits — with `await sleep(...)`, never a blocking call — until `n` tokens
  are there, then takes them; concurrent callers are served in the order they asked, and the
  bucket never over-admits.

`src/llmclient/client.py` — `AsyncLLMClient(transport, bucket, max_in_flight, *, backoff, sleep)`:

- `async ask(prompt)`: a semaphore slot, then a bucket token, then `transport.send(prompt)`;
  on `TransientError` wait `backoff`, take a **new** token and send once more; a second failure
  propagates; `asyncio.CancelledError` is never caught;
- `async ask_many(prompts)`: all prompts concurrently within the two limits, results in
  **prompt order**.

`src/llmclient/transport.py` — `MockTransport(latency=0.05, fail_every=0)` is **provided**: it
echoes the prompt after `latency` seconds, raises `TransientError` on every `fail_every`-th
call, and counts calls and in-flight sends. Do not edit it.

`clock` and `sleep` are plain parameters (defaults `time.monotonic` / `asyncio.sleep`) so the
tests can drive everything on a fake clock (`tests/conftest.py`): the whole suite runs in
milliseconds, and a real sleep anywhere is a bug the last test catches. Only the injected `sleep`
advances that clock: a bucket that waits on `asyncio.sleep` directly, or busy-loops on
`asyncio.sleep(0)`, fails the three timing tests (the busy loop after 5 s of real time, with a
message saying so, instead of hanging).

## Run it

```
cd exercises/py-07-async-token-bucket && uv sync && uv run pytest -q
```

`uv sync` creates `.venv/` with pytest, ruff and mypy (no pytest-asyncio: every test is a plain
`def` that calls `asyncio.run`). `uv run ruff check .` and `uv run mypy src` are the same bar the
reference solution clears (the untouched starter fails mypy on its `...` bodies — expected).

## The checks (`tests/test_bucket.py`, `tests/test_client.py`)

- `test_bucket_never_exceeds_rate_in_any_one_second_window` — 24 acquires at 10/s, capacity 1:
  no 1 s window holds more than 10 starts, and the last one lands after 2.3 s.
- `test_bucket_bursts_to_capacity_then_throttles` — rate 2/s, capacity 3: starts at
  0, 0, 0, 0.5, 1.0; after 10 s idle the bucket is full again, not 20 tokens deep.
- `test_ask_many_preserves_prompt_order` — replies arrive in reverse; the list comes back in
  prompt order.
- `test_in_flight_never_exceeds_max_in_flight` — 10 prompts, 3 slots: the transport sees
  exactly 3 in flight at the peak.
- `test_retries_a_transient_failure_once_with_backoff` — the failing call is retried after the
  backoff (0.35 s of virtual time in total), and a transport that always fails gets exactly two
  calls before the error propagates.
- `test_cancellation_propagates` — a task cancelled mid-send ends cancelled: no retry, no second
  send, no result.
- `test_ten_prompts_at_rate_five_take_about_two_seconds` — rate 5/s, capacity 1: ≈ 1.85 s of
  virtual time, and under 0.5 s of real time.

Done when `uv run pytest -q` prints `7 passed`, then press "mark done" on the page.
