Four generator stages — read → normalise → chunk → batch — the same shape
generator pipelines just walked. This time you write the stages, and
eight checks grade whether they are still generators once you're done: lazy, constant-memory, and
honest about the edge cases a real chunking pipeline hits.
read(source) — pull tokens one at a time. It must never turn source into a
list; a source that goes on forever (or blows up if over-consumed) has to work.normalise(tokens) — lowercase, strip, and drop anything that strips down to nothing.chunk(tokens, size, overlap) — fixed-size windows with a shared tail between consecutive
chunks, a final short chunk instead of a dropped one (but never a chunk made only of tokens the previous
one already carried), and a loud ValueError for a nonsense overlap.batch(chunks, n) — group chunks n at a time, last group short.The checks are ordinary Python and ship with the page like everything else on a static site — you could read them in devtools if you wanted. They assert what your pipeline does with real inputs (a 10,000-token stream, a 1,000,000-token one, a poisoned source that screams if you over-read it), not how you write the code to do it.
for … in source: yield …. The moment you write
list(source) or [*source] anywhere in it, you've traded laziness for a promise
you can't keep on an infinite source.token.strip().lower(), then if s: before you yield it.size, yield a copy, then slice off the
first size - overlap items before refilling — that's the shared tail. When the source runs
dry mid-fill, yield whatever's left (even if it's shorter than size) and stop — unless
what's left is only the carried-over overlap, every token of which the previous chunk already yielded:
that's a duplicate, not a chunk, so just stop. Check overlap >= size before you build
anything, or the eventual buffer never grows and you loop forever.n and once more at the end if it's non-empty.tracemalloc. If any
stage secretly builds a list of everything, this is the check that catches it — a real generator chain
never holds more than a handful of items at once, no matter how long the source is. It runs with
overlap=0 on purpose: it measures laziness alone, so a wrong stride shows up as one red row
(the arithmetic check), not two.