Packaging & serving

A trained model sitting in a notebook helps no one. Getting it to where a product can call it is a specific, well-worn path: freeze the artifact, wrap it in an API, put that in a container, and choose how requests reach it. Miss a step and you get the classic "works on my machine" failure — in production, where it costs the most.

model registryFastAPIDocker batch vs real-time/predict

From artifact to endpoint, one layer at a time

Serving a model is four wrapping steps, each solving a problem the last one left open. The trained weights become a versioned artifact; a thin API turns a function call into an HTTP endpoint; a container freezes the exact environment so it runs identically anywhere; and an orchestrator runs and scales those containers. Step outward from the model:

Wrap the model for production — ▶ play it

The step people skip is the container, and it's the one that causes the most pain. Your model depends on a specific Python, a specific PyTorch, a specific CUDA — "works on my machine" is exactly the gap a Docker image closes by shipping the environment with the code. And the API should do more than call the model: validate the input shape, handle errors, and expose a /health check so the orchestrator knows when the container is ready.

Batch or real-time? The request pattern decides

Two serving shapes, chosen by how predictions are consumed. Real-time (online) serving answers one request at a time, now — a user is waiting. Batch (offline) serving scores a big pile of inputs on a schedule and stores the results for later lookup. They have opposite cost and latency profiles, and picking the wrong one is a common, expensive mistake. Pick a use case:

How are the predictions consumed?

The tell is simple: is someone waiting for this specific prediction right now? If yes, real-time — and you care about p95 latency, autoscaling, and the latency wall. If the results are consumed later in bulk, batch — and you optimise throughput and cost per million, running big jobs on cheap, interruptible compute. Many systems do both: batch-precompute what you can, serve real-time only what must be fresh.

One request, three doors

The same FastAPI app answers five different requests four different ways, depending on which door a request walks through. /ask and /stream go all the way to the model; /healthz is answered by the router alone — it proves the process is up, not that the model is ready; /readyz checks a flag without ever calling the model. Every request is stamped with an X-Request-Id by the first middleware in the stack — pick a route and watch how far it actually travels:

Pick a route — ▶ trace it

Pick a route above.
⚠️ Traps & honesty: the pipeline here is the common shape, not the only one — serverless and managed endpoints (SageMaker, Vertex) hide the container/orchestrator steps, and tiny models may not need Kubernetes at all · "real-time vs batch" is a spectrum; micro-batching and streaming sit between them (see data architecture) · a container guarantees the same environment, not the same hardware — GPU/driver differences still matter · the batch/real-time labels here are the typical fit; a real decision also weighs cost, SLA and data freshness.
Takeaways: serving is four wraps — versioned artifact → API (validate, errors, /health) → container (freeze the environment, kill 'works on my machine') → orchestrator (run & scale) · the container is the step that prevents the most production pain · choose real-time when someone is waiting on the specific prediction (optimise latency), batch when results are consumed later in bulk (optimise throughput/cost) · combine them — precompute in batch, serve fresh in real-time · the same service answers /ask, /stream, /healthz and /readyz differently — liveness never touches the model, readiness gates it, and one stamped request id ties every log line back to the request that caused it. Next: CI/CD & deployment rolls this container out safely.

Second opinion (taught here — these corroborate): FastAPI tutorial · Docker — Get Started · MLflow Models & serving.