The GIL, threads & processes

In Java, new Thread(r).start() four times on a four-core box gives you four cores' worth of work. In CPython it gives you four threads that take turns: the interpreter holds one mutex — the Global Interpreter Lock — and only the thread holding it runs Python bytecode. The other three wait. The whole skill on this page is knowing when the interpreter lets go of that lock (sleeping, waiting on a socket, running C code such as a numpy matmul), because that is exactly when threads become worth having — and when they are not, what to reach for instead.

GILthreading multiprocessingC extensions release it

Four workloads, one timeline

Every scene below is the same shape: 8 equal units of work handed round-robin to n workers, drawn as one lane per worker with time left to right. The axis always spans the time one worker took, so the empty space to the right of the dashed wall line is what you saved. A small GIL token sits on the block that holds the interpreter; where no block holds it, the interpreter is idle or busy in C. The readout quotes wall time as a multiple of the single-worker time.

Measured once, on a 4-core box (Linux VM, Python 3.10.12, numpy 2.2.6, OpenBLAS pinned to one thread with OPENBLAS_NUM_THREADS=1); best of 3 runs. Yours will differ — the ratios are the point, not the milliseconds. The lanes are drawn from a model fitted to those wall times: the real interpreter switches threads every 5 ms (sys.getswitchinterval()), ~120 times in the CPU-bound run, not the 16 turns drawn.

What the lock actually is

CPython's memory management is reference counting, and refcounts are plain integers touched on nearly every bytecode. Rather than lock each object, CPython locks the interpreter: a thread must hold the GIL to execute bytecode, and it is asked to give it up every sys.getswitchinterval() seconds (0.005 by default) so the others get a turn. The operating system still sees four real threads on four cores; three of them are simply parked on a mutex. That is why the first scene is 1.0×: eight units of bytecode take eight units of interpreter time no matter how many threads are asking for it — and the ~120 forced hand-overs in that run cost a little, which is why threaded CPU-bound code is occasionally slower than the plain loop.

When the interpreter lets go

The GIL is released, explicitly, by any C code that is about to do something that does not touch Python objects: time.sleep, every blocking socket and file call, subprocess waits, hashlib on large buffers, zlib, and — the one that matters for this track — the BLAS call under a numpy @, which runs inside Py_BEGIN_ALLOW_THREADS. While one thread is inside such a call, another thread can hold the GIL and run bytecode. So the third scene is 4.1× with four threads and 8.0× with eight: a sleeping thread needs neither the GIL nor a core. The fourth scene is 3.6× with four threads — real parallelism from plain threading, because the matmul spends ~95% of its time in C with the lock released — but 3.3× with eight: the machine has four cores, and once the GIL is out of the way, cores are the ceiling.

One trap hides in that last scene: OpenBLAS already runs its own thread pool across your cores. Threading matmuls yourself on top of that oversubscribes the machine and can be slower than one thread. The measurement pinned BLAS to one thread (OPENBLAS_NUM_THREADS=1) to show the GIL effect in isolation; in real code you pick one level of parallelism, not two.

Processes: parallel, at a price

A process is a whole interpreter with its own GIL, so the second scene is 3.5× on four cores. The price: each process is spawned (on macOS and Windows the default start method is spawn — a fresh interpreter that re-imports your module, which is why the if __name__ == "__main__": guard is not optional there), arguments and results cross by pickling, and nothing is shared unless you set up shared memory on purpose. The one-process baseline in that scene already includes its own spawn; four processes cost ~15 ms to start on the measured Linux box and noticeably more on macOS. Cheap for a 0.6 s job; ruinous if you spawn a process per request.

The 3.13+ footnote

Since 3.13 there is a free-threaded build (python3.13t, sys._is_gil_enabled() → False) with no GIL at all — experimental in 3.13, officially supported since 3.14 (PEP 779), still not the default: CPU-bound threads scale with cores, single-threaded code runs somewhat slower, and many C-extension wheels are still not built for it. The honest interview answer in 2026 is still the picture above: threads for I/O and for C code that releases the lock, processes for pure-Python CPU work, and the free-threaded build is coming.

Which one, when

the work is…reach forwhy
pure-Python CPU (parsing, loops, regex over text)multiprocessing / concurrent.futures.ProcessPoolExecutoreach process has its own GIL; pay the spawn once, reuse the pool
waiting on network, disk, subprocessesthreading, or asyncio for many small waitsthe lock is released while blocked, so waits overlap
numpy / torch / tokenizers in C or Rustoften nothing — the library threads for youBLAS and Rust tokenizers already use your cores; adding Python threads on top oversubscribes
a mix (a service that calls an LLM and post-processes)asyncio for the calls, a process pool for the CPU parteach ceiling handled by the tool built for it
Takeaways: the GIL lets one thread run bytecode at a time, so CPU-bound Python on four threads is 1.0× — no faster, sometimes slower. The lock is released in time.sleep, blocking I/O and C calls such as a numpy matmul, so those overlap across threads (I/O even past the core count). Processes sidestep the lock entirely (3.5× on four cores here) at the cost of spawn time and pickling. Once the GIL is out of the way, cores are the next ceiling: eight matmul threads on four cores are no faster than four. Pick one level of parallelism — yours or the library's, not both.