Python's standard library is unusually good, and the parts you reach for daily are small.
Three of them earn their keep for reasons beyond terseness: Counter because
most_common(k) is a different algorithm from sorting, pathlib because paths are objects
rather than strings you glue together, and itertools because it turns a pipeline that materialises
every intermediate list into one that holds a single item at a time.
most_common is not a sortCounting is the easy half — Counter(words) against a hand-rolled dict loop is four lines
saved and no more. The interesting half is what you do next. Getting the top k by sorting everything
is O(n log n); most_common(k) keeps a heap of size k and is O(n log k).
Both are counted below by running the actual algorithms and incrementing on every comparison:
The gap widens with n and narrows as k approaches n — at most_common() with no argument
Counter sorts everything, because that is the right thing to do when you want them all. This is the general
shape of "use the library": the saving is rarely the characters, it is that somebody already picked the right
algorithm for the shape of the question you asked.
from collections import Counter, defaultdict counts = Counter(words) # the whole loop, in one call counts.most_common(3) # heap of size 3, not a full sort counts["missing"] # 0 — never a KeyError by_first = defaultdict(list) # grouping without the setdefault dance for w in words: by_first[w[0]].append(w)
Here is a three-stage pipeline over n records — filter, transform, take the first few. Written with list
comprehensions each stage builds a complete list before the next one starts; written with generators or
itertools each item is pulled through all three stages before the next item is read. Same result,
and a very different peak memory:
The second number matters as much as the memory: the lazy pipeline stops reading the source once it has
enough. If your source is a 4 GB log file, a database cursor or a paginated API, the eager version reads all
of it to build a list you then throw away. That is the whole reason
generators exist, and itertools is a library of them —
chain, islice, groupby, pairwise,
combinations, accumulate — each replacing a loop you would otherwise write with an
index variable and an off-by-one.
The last one has no measurement because the win is not a number, it is a category. os.path
treats a path as a string, so every operation is a function call wrapping the previous one, and correctness
depends on you remembering which function. pathlib makes it an object with methods and an
overloaded /:
import os, pathlib # os.path — reads inside-out, and the separator is yours to get wrong out = os.path.join(os.path.dirname(os.path.abspath(__file__)), "data", "raw.csv") name = os.path.splitext(os.path.basename(out))[0] os.makedirs(os.path.dirname(out), exist_ok=True) # pathlib — reads left to right, and / is correct on every OS out = pathlib.Path(__file__).resolve().parent / "data" / "raw.csv" name = out.stem out.parent.mkdir(parents=True, exist_ok=True) text = out.read_text() # open/read/close, in one call for csv in out.parent.glob("*.csv"): # iterate matching files directly ...
The / operator is __truediv__ on Path — the same operator
overloading you saw in dunder methods, used to make a common operation read
like the thing it means. And because Path objects know how to compare and hash, they work as dict
keys and set members, which is what you want when de-duplicating a directory walk.
groupby requires its input to be sorted by the same key, which is the single most
common itertools bug.Counter replaces the counting loop, and
most_common(k) replaces the algorithm — a size-k heap rather than a full sort · a missing
key in a Counter is 0, not a KeyError, and defaultdict(list) is the grouping idiom · an eager
pipeline materialises every stage; a generator/itertools pipeline holds one item and stops reading the source
as soon as it has enough · pathlib makes a path an object: / to join,
.stem, .parent, .read_text(), .glob() — left to right, and
correct on every OS · reach for the stdlib not to save characters but because somebody already chose the right
algorithm. Next: pytest.Second opinion (taught here — these corroborate): collections docs · itertools docs · pathlib docs.