The asyncio event loop

The GIL page showed that a thread blocked on I/O costs nothing but a thread. asyncio goes one step further: one thread, one loop, and coroutines that hand control back to the loop at every await. If you know Java's CompletableFuture chains or Loom's virtual threads, the mental model is close — with one sharp difference: nothing here is pre-emptive. A coroutine that does not await keeps the only thread, and every other task on the ring simply stops. That single fact explains the time.sleep bug, why gather is fast, and why a rate limit and a concurrency limit are two different tools.

asyncioawait gatherSemaphoretoken bucket

Five LLM calls, one loop

Every program below makes the same five calls — each one await asyncio.sleep(1.0) standing in for a 1 s request to a model — and differs only in how it lets them overlap. The clock is simulated, so every number in the readout is exact by construction: a task is ready (waiting for its turn), awaiting (its request is in flight; the arc is the timer), or done. Scrub t to watch the dots move; the strip underneath is the same story as a timeline.

Cooperative, not pre-emptive

An async def is a coroutine: calling it builds an object, it runs only when the loop drives it, and it runs until its next await on something not yet ready. At that point it parks itself on the loop with "wake me when this future resolves" and the loop picks the next ready task. There is no switch interval and no GIL hand-over — the loop is a plain while over a ready queue and a timer heap, on one thread. That is why program 2 finishes in 1.0 s: five coroutines each park on a 1 s timer, the loop sleeps once until the earliest timer, and wakes all five. And it is why program 3 takes 5.0 s again: time.sleep is a blocking syscall that never yields to the loop, so the ring stops for a full second per call. The same happens with requests.get, a synchronous database driver, or a CPU-heavy loop inside a coroutine — the fix is await asyncio.sleep, an async client (httpx.AsyncClient), or await loop.run_in_executor(None, blocking_fn) to push the work to a thread.

Two different limits

asyncio.Semaphore(2) bounds how many requests are in flight: at most two dots on the ring at once, so five calls take three rounds (3.0 s). A token bucket bounds how often you start: tokens refill at 2 per second up to a burst of 2, and a call takes one before it may begin. Nothing stops a third call being in flight while the first two are still waiting for their reply, so program 5 reaches 3 in flight and finishes in 2.5 s. An LLM provider gives you both kinds of limit (requests per minute and concurrent connections), and a client that only has one of the two will trip the other. Exercise py-07 builds both and composes them.

Cancellation is a request, delivered as an exception

task.cancel() does not kill anything. It arranges for asyncio.CancelledError to be raised inside the coroutine at its next await. Since 3.8 that is a BaseException, so an except Exception: retry loop lets it through — but a bare except:, an except BaseException:, or a "retry on anything" wrapper swallows it, and the task keeps running after its caller has given up on it. Rule: catch what you can handle, let CancelledError propagate, do your cleanup in finally.

The four calls you will actually use

callwhat it doesthe mistake it prevents
asyncio.run(main())creates the loop, runs one coroutine, closes the loophand-rolling get_event_loop
await asyncio.gather(*coros)runs them concurrently, returns results in argument ordera sequential for … await that is N× slower
async with asyncio.TaskGroup() as tg: (3.11+)gather that cancels its siblings when one failsan orphaned task that keeps running after an error
async with sem: / a bucket's await acquire()bounds in-flight count / bounds start ratea 429 storm the moment traffic arrives
Takeaways: asyncio is one thread driving coroutines that yield at await; nothing pre-empts them, so a blocking call (time.sleep, a sync HTTP client, a CPU loop) freezes every task on the ring. gather overlaps waits — five 1 s calls in 1.0 s, results in argument order. A Semaphore(n) bounds in-flight count; a token bucket bounds start rate — different limits, both needed against a real API. CancelledError is how cancellation reaches a task: let it propagate.