Boss challenge: the fraud model nobody can fool

Twelve topics on evaluation, regularization and models. Now the hard part: a rare positive class, where accuracy is a liar, the default 0.5 threshold is arbitrary, and the obvious way to cross-validate quietly cheats. Build an honest evaluation, then a model that survives it.

confusion matrix precision/recall thresholds stratified CV leakage scaling

The situation

Roughly one transaction in twenty is fraudulent. A model that predicts "never fraud" is 95% accurate and completely worthless — which is why this phase spent so long on the confusion matrix. Your job is to build the machinery that tells the truth about a classifier, and then a classifier worth telling the truth about.

Scaling before splitting is leakage. Fit a scaler on the whole dataset and every fold's "unseen" test data has already contributed its mean and variance to training. The score comes out flattering and the model disappoints in production. The scaler has to be fitted inside each fold — which is precisely what a Pipeline is for.
Plain k-fold on imbalanced data is worse than it looks. If the rows arrive sorted by class, unshuffled KFold hands you folds containing a single class. F1 against a test set with no positives is meaningless, and averaging those meaningless numbers gives you a confident, wrong estimate. Stratify.
Distance-based models cannot see past the units. With income in rupees and age in years, income is thousands of times larger, so every neighbour is chosen by income alone — the age column may as well not exist. Nothing errors; the model is simply worse for a reason that never appears in the code.

Write it

The grader runs your code, then runs 6 checks against it with data you can't see — so solving the example instead of the problem will fail. Each check reports exactly what it expected and what it got. All 6 green marks this phase ready ✓ on your roadmap.

Phase 3 · boss challenge

Where is Python coming from? No server is involved. The browser downloads CPython compiled to WebAssembly the first time you press Run & grade, along with numpy and scikit-learn, then runs your code locally. Nothing you write leaves the machine — which also means the grader is honest: it really executed what you wrote.

What passing actually proves

That you can tell a good score from an honest one. Accuracy on imbalanced data, a scaler fitted before the split, an unshuffled fold that contains one class, a threshold left at 0.5 because that's the default — each of these produces a number that looks fine and is wrong, and none of them raises. Passing means your evaluation would survive contact with a sceptical reviewer, which is the only kind worth having.

Take away: in classification the metric is the design decision. Choosing precision over recall, and the threshold that trades between them, is a statement about what a mistake costs — and nobody but you can make it.
Next: Phase 4 — the neural net playground