Feature engineering

Models don't see your data — they see the numbers you hand them. Choose those numbers badly and a good algorithm produces nonsense: one column silently drowns out every other, a category gets treated as a quantity, or a single careless line lets the answer leak into the input and your accuracy score becomes a lie.

scalingstandardisationone-hot data leakagetrain/test split

Why scale: the biggest column wins by default

Here are six loan applicants described by age (years) and income (₹ per year). We want the nearest neighbour to a new applicant — the core of kNN, and of any method built on distance (k-means, SVM with an RBF kernel, PCA).

The trouble is that income spans hundreds of thousands while age spans a few decades. In raw units, the squared difference in income is astronomically larger than anything age can contribute, so the "distance" is effectively income distance and age is ignored. Toggle scaling and watch which neighbour wins:

Nearest neighbour, before and after scaling

raw units · standardise (z-score) · min-max to [0,1]

Two standard fixes, and they are not interchangeable. Standardisation (z = (x − μ) / σ) centres each column at 0 with unit spread — the default for anything distance- or gradient-based, and it tolerates outliers reasonably. Min-max squeezes into [0, 1], which is handy when you need a bounded range (image pixels, some neural inputs) but a single extreme value compresses everyone else into a sliver. Tree models — decision trees, random forests — split on one column at a time and are therefore indifferent to scaling; don't cargo-cult a scaler in front of them.

Categories aren't numbers

Suppose the applicant also has a city: Delhi, Mumbai or Chennai. The tempting shortcut is to number them 0, 1, 2 — and that quietly tells the model Chennai (2) is twice Mumbai (1) and that Delhi and Chennai are further apart than Delhi and Mumbai. None of that is true; the labels have no order. One-hot encoding gives each category its own 0/1 column, so no false ordering is implied:

Label encoding vs one-hot — step through it

Use ordinal encoding only when the order is real (small < medium < large). Watch the column count too: one-hot on a high-cardinality column like postcode explodes your feature space — that's when you reach for target/frequency encoding or an embedding instead.

The leakage trap

This is the mistake that produces a brilliant validation score and a model that fails in production. To standardise you need a mean and a standard deviation. Compute them over the whole dataset and then split, and every training row has been nudged by information from the test set. The test set is no longer unseen, and your estimate of performance is optimistic — sometimes wildly so.

The same trap hides in imputation (filling missing values with a global mean), feature selection done before splitting, and oversampling before splitting.

Fit on train only — the order of operations

⚠️ Traps & honesty: the distances and winners here are computed live from the numbers on screen — switch scaling and the reported neighbour really is recomputed · standardising does not make data normal, it only recentres and rescales · tree ensembles don't need scaling, but the distance-based models on this page do · one-hot with all k columns is collinear for linear models (drop one, or use a regulariser) · the leakage demo shows the mechanism, not a benchmark: the size of the optimism depends on the dataset.
Takeaways: distance-based models are dominated by whichever column has the largest raw spread — scale first · standardise by default, min-max when you need a bounded range, neither for trees · one-hot unordered categories so you don't invent a fake ordering · fit every transformer on the training split only, then apply it to test — put the whole thing in a Pipeline so cross-validation can't leak. Next: cross-validation scores the pipeline you just built, and SHAP & LIME asks the fitted model to justify a single decision made from these features.

Curated companions: scikit-learn — Preprocessing · scikit-learn — Pipelines.