Twelve topics on evaluation, regularization and models. Now the hard part: a rare positive class, where accuracy is a liar, the default 0.5 threshold is arbitrary, and the obvious way to cross-validate quietly cheats. Build an honest evaluation, then a model that survives it.
Roughly one transaction in twenty is fraudulent. A model that predicts "never fraud" is 95% accurate and completely worthless — which is why this phase spent so long on the confusion matrix. Your job is to build the machinery that tells the truth about a classifier, and then a classifier worth telling the truth about.
confusion_counts — the four numbers every other metric is made of.prf — precision, recall and F1, including when they're undefined. A
model that predicts no positives has no precision; that must not be a crash.pick_threshold — 0.5 is a default, not a decision. Choose the threshold that
catches the most fraud while keeping precision above what the business will tolerate.honest_cv — cross-validation that neither leaks nor lands a fold with only
one class in it.fit_best — a model that clears the bar on data it has never seen.Pipeline is for.
KFold hands you folds containing a single class. F1 against a
test set with no positives is meaningless, and averaging those meaningless numbers gives you a
confident, wrong estimate. Stratify.
The grader runs your code, then runs 6 checks against it with data you can't see — so solving the example instead of the problem will fail. Each check reports exactly what it expected and what it got. All 6 green marks this phase ready ✓ on your roadmap.
That you can tell a good score from an honest one. Accuracy on imbalanced data, a scaler fitted before the split, an unshuffled fold that contains one class, a threshold left at 0.5 because that's the default — each of these produces a number that looks fine and is wrong, and none of them raises. Passing means your evaluation would survive contact with a sceptical reviewer, which is the only kind worth having.