Every RAG or agent service has exactly this component: a client that calls a model API under the two limits the event loop page kept apart — requests per second (a token bucket) and connections in flight (a semaphore) — retries a transient failure once, and gives up cleanly when its caller cancels. The GIL page is why the waits overlap at all. You build it against a mock transport; flagship F1 (the RAG service) swaps in a real one. This one runs on your Mac, not in the browser — it needs a real asyncio loop with tasks, locks and cancellation, which Pyodide's single-threaded loop does not exercise faithfully.
TokenBucket(rate, capacity, *, clock, sleep) — starts full; refills rate tokens
per second continuously, capped at capacity (an hour idle must not earn an hour of
burst). async acquire(n=1) waits with await sleep(…) — never a blocking call —
until n tokens are there, then takes them; concurrent callers are served in the order they
asked, and the bucket never over-admits.AsyncLLMClient(transport, bucket, max_in_flight, *, backoff, sleep) —
async ask(prompt): a semaphore slot, then a bucket token, then transport.send;
on TransientError wait backoff, take a new token and send once more; a
second failure propagates; asyncio.CancelledError is never caught.
async ask_many(prompts): all prompts concurrently within the two limits, results in
prompt order.MockTransport(latency=0.05, fail_every=0) is provided: it echoes the prompt after
latency seconds, raises TransientError on every fail_every-th
call, and counts calls and in-flight sends. Do not edit it.clock and sleep are plain parameters (defaults time.monotonic /
asyncio.sleep) so the tests can drive everything on a fake clock (tests/conftest.py):
the whole suite runs in milliseconds, and a real sleep anywhere is a bug the last test catches. Only the
injected sleep advances that clock: a bucket that waits on asyncio.sleep directly, or
busy-loops on asyncio.sleep(0), fails the three timing tests (the busy loop after 5 s of real
time, with a message saying so, instead of hanging).
The tests are ordinary pytest and ship in the public folder with the starter — read them first; the names below are the check list. Solutions are not published.
Needs git. uv installs the right Python itself, so nothing else is required.
# once, anywhere on your machine
git clone https://github.com/theDocWho/ai-ml-roadmap.git
cd ai-ml-roadmap
No git? Download the ZIP, unzip it, and cd into the unzipped folder instead.
From the repo root:
# one-time: uv (https://docs.astral.sh/uv/) manages the venv and pins Python ≥ 3.12 cd exercises/py-07-async-token-bucket && uv sync && uv run pytest -q # the same bar the reference solution clears uv run ruff check . && uv run mypy src
Done when uv run pytest -q prints 7 passed. Rerun after
every edit; pytest's -q output is the only readout this exercise has. No
pytest-asyncio: every test is a plain def that calls asyncio.run.
test_bucket_never_exceeds_rate_in_any_one_second_window — 24 acquires at 10/s, capacity 1:
no 1 s window holds more than 10 starts, and the last one lands after 2.3 s.test_bucket_bursts_to_capacity_then_throttles — rate 2/s, capacity 3: starts at
0, 0, 0, 0.5, 1.0; after 10 s idle the bucket is full again, not 20 tokens deep.test_ask_many_preserves_prompt_order — replies arrive in reverse; the list comes back in
prompt order.test_in_flight_never_exceeds_max_in_flight — 10 prompts, 3 slots: the transport sees exactly
3 in flight at the peak.test_retries_a_transient_failure_once_with_backoff — the failing call is retried after the
backoff (0.35 s of virtual time in total), and a transport that always fails gets exactly two calls
before the error propagates.test_cancellation_propagates — a task cancelled mid-send ends cancelled: no retry, no second
send, no result.test_ten_prompts_at_rate_five_take_about_two_seconds — rate 5/s, capacity 1: ≈ 1.85 s
of virtual time, and under 0.5 s of real time.This is self-attestation — the site cannot see your terminal, so the box and the button are you telling The Path the suite went green on your machine.
tokens and last; on every acquire first
add (now − last) × rate and clamp to capacity. If there is not enough, the wait
is exactly (n − tokens) / rate — one await sleep, then refill again. An
asyncio.Lock around the whole thing keeps callers in order and stops two of them spending the
same token.async with self._sem: is the concurrency bound; the bucket goes
inside it, right before the send, so a token is spent when the request actually starts.
asyncio.gather already returns results in argument order.except TransientError: is the only clause you need. Anything wider
(except Exception is fine; except BaseException or a bare except is
not) swallows the CancelledError the test sends.uv run pytest -q -x --tb=short stops at the first failure and
shows the assertion that tripped; the message names the behaviour, not the fix.