Counter, pathlib & itertools

Python's standard library is unusually good, and the parts you reach for daily are small. Three of them earn their keep for reasons beyond terseness: Counter because most_common(k) is a different algorithm from sorting, pathlib because paths are objects rather than strings you glue together, and itertools because it turns a pipeline that materialises every intermediate list into one that holds a single item at a time.

collections.Counterdefaultdict pathlib.Pathitertoolslaziness

Counter, and why most_common is not a sort

Counting is the easy half — Counter(words) against a hand-rolled dict loop is four lines saved and no more. The interesting half is what you do next. Getting the top k by sorting everything is O(n log n); most_common(k) keeps a heap of size k and is O(n log k). Both are counted below by running the actual algorithms and incrementing on every comparison:

Top-k: sort everything, or keep a heap of k?

how many keys the Counter ends up holding
most_common(k)

The gap widens with n and narrows as k approaches n — at most_common() with no argument Counter sorts everything, because that is the right thing to do when you want them all. This is the general shape of "use the library": the saving is rarely the characters, it is that somebody already picked the right algorithm for the shape of the question you asked.

from collections import Counter, defaultdict

counts = Counter(words)                # the whole loop, in one call
counts.most_common(3)                   # heap of size 3, not a full sort
counts["missing"]                       # 0 — never a KeyError

by_first = defaultdict(list)            # grouping without the setdefault dance
for w in words:
    by_first[w[0]].append(w)

itertools: the pipeline that never holds the list

Here is a three-stage pipeline over n records — filter, transform, take the first few. Written with list comprehensions each stage builds a complete list before the next one starts; written with generators or itertools each item is pulled through all three stages before the next item is read. Same result, and a very different peak memory:

Peak items held in memory

a log file, a query result, a directory walk
islice(…, take) — the first few

The second number matters as much as the memory: the lazy pipeline stops reading the source once it has enough. If your source is a 4 GB log file, a database cursor or a paginated API, the eager version reads all of it to build a list you then throw away. That is the whole reason generators exist, and itertools is a library of them — chain, islice, groupby, pairwise, combinations, accumulate — each replacing a loop you would otherwise write with an index variable and an off-by-one.

pathlib: a path is an object

The last one has no measurement because the win is not a number, it is a category. os.path treats a path as a string, so every operation is a function call wrapping the previous one, and correctness depends on you remembering which function. pathlib makes it an object with methods and an overloaded /:

import os, pathlib

# os.path — reads inside-out, and the separator is yours to get wrong
out = os.path.join(os.path.dirname(os.path.abspath(__file__)), "data", "raw.csv")
name = os.path.splitext(os.path.basename(out))[0]
os.makedirs(os.path.dirname(out), exist_ok=True)

# pathlib — reads left to right, and / is correct on every OS
out = pathlib.Path(__file__).resolve().parent / "data" / "raw.csv"
name = out.stem
out.parent.mkdir(parents=True, exist_ok=True)
text = out.read_text()                  # open/read/close, in one call
for csv in out.parent.glob("*.csv"):    # iterate matching files directly
    ...

The / operator is __truediv__ on Path — the same operator overloading you saw in dunder methods, used to make a common operation read like the thing it means. And because Path objects know how to compare and hash, they work as dict keys and set members, which is what you want when de-duplicating a directory walk.

⚠️ Traps & honesty: the comparison counts come from running a merge sort and a size-k heap in this page — CPython's Timsort exploits existing order and beats the merge-sort count on partially sorted input, so treat the sort column as an upper bound · both algorithms are fast enough that on a few thousand keys you will not feel the difference; the point is that the library picked the better one for free · the memory figure counts items held, not bytes, and ignores the interpreter's own overhead · laziness is not free: a generator can only be consumed once, and debugging one is harder because there is no list to print · groupby requires its input to be sorted by the same key, which is the single most common itertools bug.
Takeaways: Counter replaces the counting loop, and most_common(k) replaces the algorithm — a size-k heap rather than a full sort · a missing key in a Counter is 0, not a KeyError, and defaultdict(list) is the grouping idiom · an eager pipeline materialises every stage; a generator/itertools pipeline holds one item and stops reading the source as soon as it has enough · pathlib makes a path an object: / to join, .stem, .parent, .read_text(), .glob() — left to right, and correct on every OS · reach for the stdlib not to save characters but because somebody already chose the right algorithm. Next: pytest.

Second opinion (taught here — these corroborate): collections docs · itertools docs · pathlib docs.