ML system design

"Design a video recommender." The interview question that sinks strong engineers isn't hard because of the model — it's that they dive straight into "I'd use a transformer" and skip everything that actually matters. The people who pass follow a framework: a fixed running order that starts with the problem, not the model. Learn the order and the question stops being scary.

requirementsdata & scalefeatures metricsservingmonitoring

The running order — start with the problem, not the model

Every good answer walks the same seven stations, in order. The model is one of them, and it's in the middle — the stations before it decide whether you're even solving the right problem, and the ones after decide whether it survives contact with reality. Step through the framework:

The 7-station framework — ▶ play it

The two stations candidates skip are the first and the last. Clarify requirements first — scale, latency budget, and above all what "good" means as a metric — because "recommend videos" could mean maximise watch time, diversity, or new-creator discovery, and they lead to different systems. And monitoring last, because a model that isn't watched for drift silently rots. Say those two out loud and you're already ahead of most candidates.

Same framework, any problem

The power of a framework is that it transfers. Pick a different "design X" and the seven stations stay put — only the answers change. See the whole design fill in for a concrete problem:

Pick a problem, see the design fill in

Notice what changes between problems and what doesn't. Fraud is extreme class imbalance and a hard latency budget (score before the transaction clears), so precision/recall trade-offs and real-time serving dominate. Recommendation is a two-stage candidate-generation-then-ranking design at massive scale. Search is learning-to-rank with relevance as the metric. Different answers, identical stations — that's the whole point of having a framework.

⚠️ Traps & honesty: this is a framework, not a script — a real interview is a conversation where you go deep where the interviewer pushes, not a monologue through seven boxes · the worked designs here are one reasonable sketch each, deliberately compressed; real answers dive much deeper on one or two stations · the biggest mistake isn't a wrong model, it's jumping to the model before clarifying the metric and the constraints · "scale" numbers should be estimated live (back-of-envelope), not memorised · a good answer states assumptions and trade-offs out loud rather than presenting one true design.
Takeaways: answer "design X" with a fixed running order — requirements → data & scale → features → model → evaluation → serving → monitoring — and start with the problem, not the model · clarify what "good" means as a metric before anything else; the same prompt can mean different systems · the model is one middle station, not the whole answer · the stations transfer to any problem — only the answers change · state assumptions and trade-offs out loud. This closes the roadmap: from Python to a system you can design and defend.

Second opinion (taught here — these corroborate): Chip Huyen — ML Systems Design · System Design Primer.