Generators & yield showed that a single
yield remembers its place and a list pays for everything up front. Real pipelines chain
several generators together — read → normalise → chunk → batch — and that's where
the interesting failures live: what "one item in flight" really means across four stages, why draining
a generator twice is a silent no-op, and the three itertools tools you reach for weekly.
Each box below is a generator function; the pipe between two boxes is one generator handing values to the
next. Pull a value from batch and it calls next() on chunk, which
calls next() on normalise, which calls next() on read —
the demand runs backwards through the chain, one item at a time, and nothing downstream moves until
something upstream is asked for. Pick an operation and watch each stage's own counter — "items passed" — the token on
the pipe between two stages, and the memory meter underneath, which compares this lazy chain against
materialising a list at every stage.
items resident at once at this N — the lazy chain against a list materialised at every stage:
for hides itA generator is a cursor over a computation, not a container: once it has yielded its last value, it
stays exhausted forever. Calling next() on a spent generator doesn't restart it — it raises
StopIteration, every time, for the rest of that object's life. A for loop hides
this completely, because a for loop's entire job is to call next() until it sees
exactly that exception and then stop quietly — which is also why for item in gen: ... followed
by a second for item in gen: ... silently does nothing the second time. If you need two passes,
you need two generators (call the generator function again), or materialise once with
list(gen) and accept the memory cost. Java's Stream has the same rule — reuse one
after a terminal operation and you get IllegalStateException: stream has already been operated upon
or closed — the difference is that Python's version fails quietly: in a pipeline, a stage
handed an already-drained generator just produces nothing, and the bug shows up downstream as an empty
batch, not as an exception at the point of reuse.
send and yield as a two-way channelyield is usually read as one-directional — the generator hands a value out — but it is
actually an expression with a value of its own: x = yield produced pauses after handing out
produced, and when the caller resumes it with gen.send(v), that call becomes the
value of the yield expression, so x is now v inside the
generator. The arrow of data runs backwards, into a stage that is already paused mid-flight — a live channel
into a running pipeline stage, not just a queue you read from.
islice, chain, tee — the three you use weeklyitertools.islice(it, n) — take only the first n items without
materialising anything before them, and without asking the source for more than it has to. A downstream
buffer (like a chunker holding one item of lookahead) can make the source's own counter read slightly
higher than n — that lookahead item was pulled but never yielded onward.itertools.chain(a, b, ...) — walk several iterables as if they were one, lazily:
it does not concatenate anything, it just keeps asking the current source for next() and
moves to the next source only once the current one raises StopIteration.itertools.tee(it, n) — split one iterator into n independent ones. The
source is still read exactly once; whichever copy has consumed fewer items forces tee to
buffer the difference internally, so two consumers running at very different speeds cost memory
proportional to how far apart they've drifted, not to the whole source.for loop's silence about that is a trap, not a feature. send() turns
yield into a two-way channel into a paused stage. islice stops the source early,
chain walks several sources as one, and tee buffers only the gap between two
consumers of one source.