The GIL page showed that a thread blocked on I/O
costs nothing but a thread. asyncio goes one step further: one thread, one loop, and coroutines
that hand control back to the loop at every await. If you know Java's
CompletableFuture chains or Loom's virtual threads, the mental model is close — with one
sharp difference: nothing here is pre-emptive. A coroutine that does not await keeps the
only thread, and every other task on the ring simply stops. That single fact explains the
time.sleep bug, why gather is fast, and why a rate limit and a concurrency
limit are two different tools.
Every program below makes the same five calls — each one await asyncio.sleep(1.0) standing
in for a 1 s request to a model — and differs only in how it lets them overlap. The clock is
simulated, so every number in the readout is exact by construction: a task is ready (waiting for
its turn), awaiting (its request is in flight; the arc is the timer), or done. Scrub
t to watch the dots move; the strip underneath is the same story as a timeline.
An async def is a coroutine: calling it builds an object, it runs only when the loop drives
it, and it runs until its next await on something not yet ready. At that point it
parks itself on the loop with "wake me when this future resolves" and the loop picks the next ready task.
There is no switch interval and no GIL hand-over — the loop is a plain while over a ready
queue and a timer heap, on one thread. That is why program 2 finishes in 1.0 s: five
coroutines each park on a 1 s timer, the loop sleeps once until the earliest timer, and wakes all
five. And it is why program 3 takes 5.0 s again: time.sleep is a blocking
syscall that never yields to the loop, so the ring stops for a full second per call. The same happens with
requests.get, a synchronous database driver, or a CPU-heavy loop inside a coroutine — the
fix is await asyncio.sleep, an async client (httpx.AsyncClient), or
await loop.run_in_executor(None, blocking_fn) to push the work to a thread.
asyncio.Semaphore(2) bounds how many requests are in flight: at most two dots on the
ring at once, so five calls take three rounds (3.0 s). A token bucket bounds how often
you start: tokens refill at 2 per second up to a burst of 2, and a call takes one before it may begin.
Nothing stops a third call being in flight while the first two are still waiting for their reply, so
program 5 reaches 3 in flight and finishes in 2.5 s. An LLM provider gives you both
kinds of limit (requests per minute and concurrent connections), and a client that only has one
of the two will trip the other. Exercise py-07 builds both and composes them.
task.cancel() does not kill anything. It arranges for asyncio.CancelledError to
be raised inside the coroutine at its next await. Since 3.8 that is a BaseException,
so an except Exception: retry loop lets it through — but a bare except:, an
except BaseException:, or a "retry on anything" wrapper swallows it, and the task keeps
running after its caller has given up on it. Rule: catch what you can handle, let CancelledError
propagate, do your cleanup in finally.
| call | what it does | the mistake it prevents |
|---|---|---|
asyncio.run(main()) | creates the loop, runs one coroutine, closes the loop | hand-rolling get_event_loop |
await asyncio.gather(*coros) | runs them concurrently, returns results in argument order | a sequential for … await that is N× slower |
async with asyncio.TaskGroup() as tg: (3.11+) | gather that cancels its siblings when one fails | an orphaned task that keeps running after an error |
async with sem: / a bucket's await acquire() | bounds in-flight count / bounds start rate | a 429 storm the moment traffic arrives |
await; nothing pre-empts them, so a blocking call (time.sleep, a sync HTTP
client, a CPU loop) freezes every task on the ring. gather overlaps waits — five 1 s
calls in 1.0 s, results in argument order. A Semaphore(n) bounds in-flight
count; a token bucket bounds start rate — different limits, both needed against a real API.
CancelledError is how cancellation reaches a task: let it propagate.