Models don't see your data — they see the numbers you hand them. Choose those numbers badly and a good algorithm produces nonsense: one column silently drowns out every other, a category gets treated as a quantity, or a single careless line lets the answer leak into the input and your accuracy score becomes a lie.
Here are six loan applicants described by age (years) and income (₹ per year). We want the nearest neighbour to a new applicant — the core of kNN, and of any method built on distance (k-means, SVM with an RBF kernel, PCA).
The trouble is that income spans hundreds of thousands while age spans a few decades. In raw units, the squared difference in income is astronomically larger than anything age can contribute, so the "distance" is effectively income distance and age is ignored. Toggle scaling and watch which neighbour wins:
Two standard fixes, and they are not interchangeable. Standardisation (z = (x − μ) / σ)
centres each column at 0 with unit spread — the default for anything distance- or gradient-based, and it
tolerates outliers reasonably. Min-max squeezes into [0, 1], which is handy when you need
a bounded range (image pixels, some neural inputs) but a single extreme value compresses everyone else into a
sliver. Tree models — decision trees,
random forests — split on one column at a time and are therefore
indifferent to scaling; don't cargo-cult a scaler in front of them.
Suppose the applicant also has a city: Delhi, Mumbai or Chennai. The tempting shortcut is to number them 0, 1, 2 — and that quietly tells the model Chennai (2) is twice Mumbai (1) and that Delhi and Chennai are further apart than Delhi and Mumbai. None of that is true; the labels have no order. One-hot encoding gives each category its own 0/1 column, so no false ordering is implied:
Use ordinal encoding only when the order is real (small < medium < large). Watch the column count too: one-hot on a high-cardinality column like postcode explodes your feature space — that's when you reach for target/frequency encoding or an embedding instead.
This is the mistake that produces a brilliant validation score and a model that fails in production. To standardise you need a mean and a standard deviation. Compute them over the whole dataset and then split, and every training row has been nudged by information from the test set. The test set is no longer unseen, and your estimate of performance is optimistic — sometimes wildly so.
The same trap hides in imputation (filling missing values with a global mean), feature selection done before splitting, and oversampling before splitting.
Pipeline so
cross-validation can't leak. Next: cross-validation scores the pipeline
you just built, and SHAP & LIME asks the fitted model to justify a single
decision made from these features.Curated companions: scikit-learn — Preprocessing · scikit-learn — Pipelines.