CI/CD & deployment

The riskiest moment in a model's life is the instant it goes live. Ship the new version to 100% of traffic and pray, and a bad model becomes a company-wide incident in seconds. Every safe deployment strategy is a way to limit the blast radius — to expose the new version to a little reality first, watch, and be able to undo it instantly.

canaryblue-greenshadow rollbackCI/CD

The canary: expose a little, watch, then commit

A canary release sends a small slice of live traffic to the new model while everyone else stays on the old one. You watch the new version's error rate against a budget; if it stays healthy you raise the slice step by step, and if it breaches the budget you roll back — having harmed only that slice, not everyone. Turn up the traffic and see whether this canary survives:

Roll out the canary — how much traffic, and does it hold?

you don't know this in advance — the canary is how you find out
the aggregate error rate you're willing to tolerate

The canary's whole value is the small first slice. At 5% traffic, a broken new model degrades 5% of requests for a few minutes before your monitoring trips and rolls back — an annoyance, not an outage. The old version never went away, so rollback is instant: shift the slice back to zero. This is why you deploy in steps (5% → 25% → 50% → 100%) with a health check between each, never in one leap.

Three strategies, three trade-offs

The canary is one point on a spectrum. Each strategy trades cost, risk and realism differently — pick by how much you can spend and how badly a bad release would hurt:

Compare the rollout strategies

Under all of them sits CI/CD: continuous integration runs your tests, data checks and model evaluation on every change, and continuous deployment promotes an artifact through the stages automatically only if the gates pass. The deployment strategy is the last gate — the one that limits damage when a bad model slips past all the earlier ones, because eventually one will.

⚠️ Traps & honesty: the aggregate-error formula here is the simple weighted blend (old rate on old traffic + new rate on new traffic) to show the blast-radius idea — real monitoring watches latency, business metrics and per-segment errors too · a canary needs enough traffic in the slice to detect a problem statistically; 1% of low volume tells you nothing · blue-green doubles your serving cost during the overlap · shadow testing must never let the shadow model's outputs reach users or cause side effects (double writes, duplicate emails) · rollback is only instant if the old version is still running and state is compatible.
Takeaways: deploying to 100% at once makes a bad model an instant outage — every safe strategy limits the blast radius · canary: a small live slice, watch against an error budget, step up or roll back · blue-green: a full standby you switch to instantly (and back), at double cost · shadow: the new model sees real traffic but affects nobody — the safest test, no user risk · CI/CD automates the gates, and the rollout strategy is the last one. Next: the MLOps lifecycle closes the loop with monitoring and drift.

Second opinion (taught here — these corroborate): MLOps Zoomcamp — deployment · Fowler — Canary release · Fowler — Blue-green.