In Java, new Thread(r).start() four times on a four-core box gives you four
cores' worth of work. In CPython it gives you four threads that take turns: the interpreter holds
one mutex — the Global Interpreter Lock — and only the thread holding it runs Python bytecode. The
other three wait. The whole skill on this page is knowing when the interpreter lets go of that
lock (sleeping, waiting on a socket, running C code such as a numpy matmul), because that is exactly when
threads become worth having — and when they are not, what to reach for instead.
Every scene below is the same shape: 8 equal units of work handed round-robin to n
workers, drawn as one lane per worker with time left to right. The axis always spans the time
one worker took, so the empty space to the right of the dashed wall line is what you saved.
A small GIL token sits on the block that holds the interpreter; where no block holds it, the
interpreter is idle or busy in C. The readout quotes wall time as a multiple of the single-worker time.
Measured once, on a 4-core box (Linux VM, Python 3.10.12,
numpy 2.2.6, OpenBLAS pinned to one thread with OPENBLAS_NUM_THREADS=1); best of 3 runs.
Yours will differ — the ratios are the point, not the milliseconds. The lanes are drawn from a model
fitted to those wall times: the real interpreter switches threads every 5 ms
(sys.getswitchinterval()), ~120 times in the CPU-bound run, not the 16 turns drawn.
CPython's memory management is reference counting, and refcounts are plain integers touched on nearly
every bytecode. Rather than lock each object, CPython locks the interpreter: a thread must hold
the GIL to execute bytecode, and it is asked to give it up every sys.getswitchinterval()
seconds (0.005 by default) so the others get a turn. The operating system still sees four real threads on
four cores; three of them are simply parked on a mutex. That is why the first scene is 1.0×: eight
units of bytecode take eight units of interpreter time no matter how many threads are asking for it —
and the ~120 forced hand-overs in that run cost a little, which is why threaded CPU-bound code is
occasionally slower than the plain loop.
The GIL is released, explicitly, by any C code that is about to do something that does not touch Python
objects: time.sleep, every blocking socket and file call, subprocess waits,
hashlib on large buffers, zlib, and — the one that matters for this track — the
BLAS call under a numpy @, which runs inside Py_BEGIN_ALLOW_THREADS. While one
thread is inside such a call, another thread can hold the GIL and run bytecode. So the third scene is
4.1× with four threads and 8.0× with eight: a sleeping thread needs neither the GIL nor a
core. The fourth scene is 3.6× with four threads — real parallelism from plain threading,
because the matmul spends ~95% of its time in C with the lock released — but 3.3× with eight:
the machine has four cores, and once the GIL is out of the way, cores are the ceiling.
One trap hides in that last scene: OpenBLAS already runs its own thread pool across your cores.
Threading matmuls yourself on top of that oversubscribes the machine and can be slower than one thread.
The measurement pinned BLAS to one thread (OPENBLAS_NUM_THREADS=1) to show the GIL effect
in isolation; in real code you pick one level of parallelism, not two.
A process is a whole interpreter with its own GIL, so the second scene is 3.5× on four cores.
The price: each process is spawned (on macOS and Windows the default start method is spawn —
a fresh interpreter that re-imports your module, which is why the if __name__ == "__main__":
guard is not optional there), arguments and results cross by pickling, and nothing is shared unless
you set up shared memory on purpose. The one-process baseline in that scene already includes its own
spawn; four processes cost ~15 ms to start on the measured Linux box and noticeably more on macOS.
Cheap for a 0.6 s job; ruinous if you spawn a process per request.
Since 3.13 there is a free-threaded build (python3.13t,
sys._is_gil_enabled() → False) with no GIL at all — experimental in 3.13, officially
supported since 3.14 (PEP 779), still not the default: CPU-bound threads scale with cores, single-threaded
code runs somewhat slower, and many C-extension wheels are still not built for it.
The honest interview answer in 2026 is still the picture above: threads for I/O and for C code that
releases the lock, processes for pure-Python CPU work, and the free-threaded build is coming.
| the work is… | reach for | why |
|---|---|---|
| pure-Python CPU (parsing, loops, regex over text) | multiprocessing / concurrent.futures.ProcessPoolExecutor | each process has its own GIL; pay the spawn once, reuse the pool |
| waiting on network, disk, subprocesses | threading, or asyncio for many small waits | the lock is released while blocked, so waits overlap |
| numpy / torch / tokenizers in C or Rust | often nothing — the library threads for you | BLAS and Rust tokenizers already use your cores; adding Python threads on top oversubscribes |
| a mix (a service that calls an LLM and post-processes) | asyncio for the calls, a process pool for the CPU part | each ceiling handled by the tool built for it |
time.sleep, blocking I/O and C calls such as a numpy matmul, so those overlap across threads
(I/O even past the core count). Processes sidestep the lock entirely (3.5× on four cores here) at the cost
of spawn time and pickling. Once the GIL is out of the way, cores are the next ceiling: eight
matmul threads on four cores are no faster than four. Pick one level of parallelism — yours or the
library's, not both.