A trained model sitting in a notebook helps no one. Getting it to where a product can call it is a specific, well-worn path: freeze the artifact, wrap it in an API, put that in a container, and choose how requests reach it. Miss a step and you get the classic "works on my machine" failure — in production, where it costs the most.
Serving a model is four wrapping steps, each solving a problem the last one left open. The trained weights become a versioned artifact; a thin API turns a function call into an HTTP endpoint; a container freezes the exact environment so it runs identically anywhere; and an orchestrator runs and scales those containers. Step outward from the model:
The step people skip is the container, and it's the one that causes the most pain. Your model depends
on a specific Python, a specific PyTorch, a specific CUDA — "works on my machine" is exactly the gap a Docker
image closes by shipping the environment with the code. And the API should do more than call the
model: validate the input shape, handle errors, and expose a /health check so the orchestrator
knows when the container is ready.
Two serving shapes, chosen by how predictions are consumed. Real-time (online) serving answers one request at a time, now — a user is waiting. Batch (offline) serving scores a big pile of inputs on a schedule and stores the results for later lookup. They have opposite cost and latency profiles, and picking the wrong one is a common, expensive mistake. Pick a use case:
The tell is simple: is someone waiting for this specific prediction right now? If yes, real-time — and you care about p95 latency, autoscaling, and the latency wall. If the results are consumed later in bulk, batch — and you optimise throughput and cost per million, running big jobs on cheap, interruptible compute. Many systems do both: batch-precompute what you can, serve real-time only what must be fresh.
The same FastAPI app answers five different requests four different ways, depending on which
door a request walks through. /ask and /stream go all the way to the
model; /healthz is answered by the router alone — it proves the process is up, not
that the model is ready; /readyz checks a flag without ever calling the model. Every
request is stamped with an X-Request-Id by the first middleware in the stack — pick
a route and watch how far it actually travels:
/ask, /stream, /healthz and /readyz
differently — liveness never touches the model, readiness gates it, and one stamped request id ties
every log line back to the request that caused it. Next:
CI/CD & deployment rolls this container out safely.Second opinion (taught here — these corroborate): FastAPI tutorial · Docker — Get Started · MLflow Models & serving.