Data architecture

A model is only as good as the data plumbing behind it. Three decisions decide whether that plumbing helps or quietly poisons your model: how fresh the data must be, whether training and serving compute features the same way, and where the data actually lives. Get the middle one wrong and your model scores well offline and fails in production — for reasons no metric will show you.

batch vs streamingfeature store train/serve skewlakehousepipelines

Freshness decides the pipeline

"Batch or streaming?" isn't a taste question — it's set by how stale your data is allowed to be. A daily sales forecast is fine on yesterday's data; a fraud check on a live transaction is not. Freshness costs money and complexity, so you buy exactly as much as the use case needs and no more. Drag the freshness requirement and watch the right architecture — and its cost — change:

How fresh must the data be?

how old the data feeding a prediction is allowed to be

The three regimes: batch (recompute on a schedule — cheapest, simplest, hours-to-days stale), micro-batch (every few minutes — a middle ground), and streaming (process each event as it arrives — seconds fresh, but you now run Kafka-style infrastructure and reason about out-of-order and late-arriving events). Most organisations over-reach for streaming; start at the cheapest regime the use case tolerates.

The bug a feature store exists to prevent

Here is the most expensive subtle bug in production ML. At training time you compute a feature — say "customer's average spend over 30 days" — in a notebook, over the full history, in Pandas. At serving time a different service, in a different language, recomputes it live and slightly differently: a different window, a timezone, a rounding rule. The model was trained on one definition and is served another. That gap is train/serve skew, and it degrades predictions silently — offline metrics look fine because they use the training definition.

A feature store fixes it by making both paths read the same feature definition. Toggle it and watch the skew appear and vanish:

Train/serve skew, with and without a feature store

off = two teams reimplement the feature · on = one shared definition
only matters when the store is OFF

Beyond consistency, a feature store gives you reuse (compute "30-day spend" once, many models read it) and point-in-time correctness (when building training data, join each label to the feature values as they were at that moment, never leaking the future — the same leakage trap from feature engineering, now at pipeline scale).

Where the data lives: lake, warehouse, or lakehouse

Three storage shapes, one trade-off between flexibility and structure. Pick a workload and see which fits:

Pick a workload, get the store

⚠️ Traps & honesty: the freshness bands and costs here are illustrative categories, not benchmarks — your real numbers depend on volume and tooling · the skew number is a simple model of "two implementations diverge", to make a real, invisible failure visible; production skew is measured by comparing logged serving features against recomputed training features · a feature store is not free infrastructure — for one model and one team, a shared library may be enough · lakehouse is a real convergence, but "just use a lakehouse" still needs governance to not become a swamp.
Takeaways: pick batch / micro-batch / streaming from the freshness the use case actually needs — streaming is powerful and over-chosen · train/serve skew is a silent killer: the same feature computed two ways degrades production quietly, and a feature store's job is one shared definition (plus reuse and point-in-time-correct joins) · match storage to the workload — lake for raw/flexible, warehouse for governed SQL, lakehouse when you need both. Next: system architecture assembles these into a design you can draw.

Second opinion (taught here — these corroborate): Data Engineering Zoomcamp · Feast — feature store docs · Databricks — Lakehouse.