All insights
2 Oct 2026 · 4 min readBy the VentureSEA Digital engineering team

What It Actually Takes to Run Machine Learning in Production

The model that earned applause in the demo meeting is maybe a fifth of the system that keeps it true in month nine. A field guide to the five capabilities that separate production ML from an expensive notebook.

What It Actually Takes to Run Machine Learning in Production

Every production ML story starts in the same meeting. The model hits its accuracy target on the holdout set, the demo lands, everyone agrees to ship it. Six months later one of two things is true: the model still is not in production, or it is, and nobody can say whether its predictions are still any good. We have taken systems through this gap for credit scoring, building controls, and insurance quality assurance, and the consistent lesson is that the model is the smallest part of the work. The rest is an operating system around it, and that operating system, not the model, is what a buyer is actually procuring when they buy production ML.

A model is a result. Production ML is a capability. Buying tools gets you neither.

Why good models die between the notebook and production

  • No promotion path: deployment happens by a person copying artefacts, so every release is bespoke, unrepeatable, and impossible to roll back cleanly.
  • Training-serving skew: the features computed in the notebook and the features computed live quietly diverge, and offline accuracy stops describing the system users actually experience.
  • Silent drift: the world moves, inputs shift, and without monitoring the first sign of decay is a business metric that has already been wrong for a quarter.
  • No retraining loop: even a monitored model just ages visibly instead of invisibly unless fresh data can flow back into a governed retrain.
  • No audit trail: in a regulated setting, a prediction nobody can reconstruct is a liability, and the system that cannot answer which model, which data, who approved does not get a production sign-off at all.

The five capabilities, in the order they pay

Everything a production ML system needs collapses into five capabilities, and the sequence matters more than the tooling:

  • Versioned promotion with gates: every deployed model is tied to a version, a training-data snapshot, and a named approver, and promotion is blocked unless the evaluation suite passes. A registry that cannot refuse a model is a filing cabinet.
  • Monitoring that pages someone: input drift and live metric decay tracked continuously, with alerts wired to a person. We have written before about how a monitored threshold is a maintained artefact, not a launch setting; the same discipline applies to every model in production.
  • A governed retraining loop: a pipeline that can rebuild the model on fresh data, re-run the gates, and promote through the same door as the original, triggered by measured drift rather than calendar habit.
  • An audit trail per inference: input, model version, output, and any human review in an append-only record. This is the governance stack seen from the MLOps side, and in finance, insurance, and government it is the difference between shipping and staying a proof of concept.
  • Human review where stakes demand it: low-confidence and high-impact predictions routed to people, with the queue sized to real volume so review is a control, not a bottleneck.

Platform shopping is the expensive way to avoid the question

The tooling for all five capabilities is commodity: MLflow for the registry, Airflow for pipelines, standard monitoring stacks, Kubernetes underneath. Teams stall anyway, because the hard part is not installing tools but wiring them to your specific risks: deciding what the eval gate must refuse, which drift threshold pages whom, what the audit record must prove to your regulator. A platform bought before those decisions exist becomes a dashboard nobody is paged on. The same discipline governs serving economics: whether inference belongs on a hosted API or your own GPUs is a utilisation question with a calculable break-even, not a platform feature, and it changes quarter by quarter.

What this looks like when it works

In a building-systems engagement, the fault-detection model was a few months of work; the operating system around it is why it still earns trust: a 93 percent true-positive rate held at roughly 7 percent false positives through seasonal threshold reviews, drift-triggered retraining, and alert suppression tuned with the operations team, contributing to 15 to 20 percent energy savings that are measurable at the meter. In a financial-services engagement, the credit model mattered less than the rails around it: registry promotion with backtesting gates and drift monitoring is what let a regulated lender cut non-performing loans by roughly a fifth while keeping acceptance rates, because the model could be retrained and redeployed with evidence instead of argument. The pattern repeats: the model earns the headline number, the operating system keeps it true.

A build sequence that avoids the platform trap

  1. Take one model end to end on thin rails: minimal registry, one eval gate that can actually refuse promotion, one drift monitor that pages a named person. Weeks, not quarters, and consistent with a 6 to 12 week first production path.
  2. Add the audit trail before adding the second model: retrofitting provenance across a fleet is the expensive version of a cheap early decision.
  3. Build the retraining loop when drift data justifies it, using the evidence the monitor has been collecting since step one.
  4. Onboard the second and third model onto the same rails. This is where the capability starts compounding: each new model inherits governance instead of re-litigating it.
  5. Only now evaluate platform purchases, against the habits you have proven, so the platform automates what works instead of standing in for what does not exist.

The notebook-to-production gap is not closed by a bigger model or a platform invoice. It is closed by five capabilities built in the right order, each one enforced in the pipeline rather than described in a slide. That is also how we structure AI consulting engagements that end in running systems: the model is scoped in weeks, the operating system is the engagement, and the handover is a capability your team runs, not a dashboard you rent.

Read the case study

AI Suite Enablement for an HVAC BMS Platform

A modular AI suite for a building management platform: fault detection, remaining-useful-life prediction, and energy optimisation.

See how we built it

Have a topic you want us to cover?

Reach out with the challenge you are working on. We write about what matters in production.