k-Nearest Neighbors

The simplest classifier there is: to label a new point, look at the k closest training points and take a vote. There's no "training" — the data is the model. Its one knob, k, is a perfect, visible illustration of the bias–variance trade-off.

distancemajority vote k = bias/variance knobnon-parametric decision boundary

Classify by neighbourhood

kNN makes a prediction with one rule: measure the distance from the new point to every training point, keep the k nearest, and let them vote — the majority class wins. There's no equation to fit; you just store the data and search it at prediction time (which is why it's called a lazy, non-parametric method). Click anywhere to drop a query point and watch its k neighbours get picked and vote:

click the canvas to move the query ●

k is the bias–variance dial

Look at the coloured background — that's the decision boundary, the region where each class would win:

Slide k from 1 upward and watch the boundary go from jagged to smooth — the same over-fit → under-fit journey you see in every model, made visible. The sweet spot is found with cross-validation.

Two practical catches

kNN relies entirely on distance, so it has two well-known weaknesses. First, features must be scaled — if one axis is "salary" (0–100,000) and another is "age" (0–100), salary dominates the distance and age is ignored; standardise first. Second, it suffers in high dimensions (the "curse of dimensionality": everything becomes roughly equidistant) and is slow at prediction time on big datasets, since every query scans the whole training set.

Takeaways: kNN labels a point by a majority vote of its k nearest neighbours — no training, the data is the model. k is the bias–variance knob: small k = jagged, overfit (low bias/high variance); large k = smooth, underfit. Scale features first (distance-based), and mind the cost in high dimensions.