MLOps for FDEs: five things to build before a model runs in a client's system
The first time a client calls to say the model's forecasts are wrong, you need immediate answers to three questions: which model is running, when it started going wrong, and whether the fault lies in the data or the model.

In brief
- The hard part of ML in production is building and running the whole system; the model is only one part of it.
- Version models as artifacts, use a single feature source, measure skew and test the pipeline with a simple model first.
- Palantir's job posting for a Forward Deployed Infrastructure Engineer states plainly that monitoring, alerting and handling production incidents, including on-call, are part of the role.
Imagine you have just put a demand-forecasting model into a supermarket chain’s ordering system. In the notebook the error looked excellent, and the demo for senior management went smoothly. Three weeks later the warehouse manager calls: orders for fresh milk have doubled, and nobody in the room knows which version of the model is running.
This situation is rarely caused by a bad model. Google Cloud’s MLOps architecture documentation says it plainly: the hard part is building an integrated ML system and keeping it running continuously in production. At least one company hiring FDEs writes this work straight into the job description.
Palantir’s posting for a Forward Deployed Infrastructure Engineer lists the work to be done at the deployment site: monitoring and alerting, configuration management, system upgrades.
The person hired also shares responsibility for diagnosing, resolving and preventing production incidents, including on-call duty. If you are aiming for roles like this, these are skills tested in daily work, not lines to dress up a CV.
Why isn’t the model the product?
In 2015, D. Sculley’s team published “Hidden Technical Debt in Machine Learning Systems” at NIPS. They observed that real-world ML systems often carry large, long-running maintenance costs, from sources such as data dependencies and changes in the outside world.
The supermarket changes its product codes, the POS terminals get a software update, shoppers change their habits after Tết, the Lunar New Year. The model stays the same, but the data it receives is no longer the same.
ML therefore has an extra requirement that ordinary software does not. Google Cloud calls it continuous training (CT): automatically retraining and then serving the new model.
They also note that CD in ML is no longer about deploying a software package or a service, but about deploying an entire system, and that CI must test the data as well as each component.
Martin Zinkevich, in “Rules of Machine Learning”, offers a short piece of advice: keep the first model simple and get the infrastructure right. For an FDE, the priorities are therefore fairly clear. The first days at a client site should go into building the data pipeline, not tuning parameters.
A worked example with the supermarket chain: five things, in order
First, know exactly what is running. Go back to the call about fresh milk: the first question is always “which model?”. The CD4ML article on Martin Fowler’s site proposes treating the model as an artifact, versioned and deployed like any other build. At a minimum, every prediction log line must record enough to trace it back:
log_prediction({
"model_version": "demand-v14",
"feature_snapshot": "2026-10-01",
"store_id": "HCM-021",
"sku": "SUA-TUOI-1L",
"prediction": 240,
})
With this log line you can answer in minutes rather than days. You can also roll back to the previous version while the warehouse is still waiting for orders.
Next, use only one feature source. Suppose the feature “7-day average sales” is computed at training time from the data warehouse, after returns have been deducted. In production, the same feature comes from the real-time POS stream, where returns have not yet been deducted. If a product nets 100 units a day in the warehouse but the POS reports 115, the model sees a world 15% busier than the one it learned from. That is training-serving skew.
Google Cloud recommends a feature store as the shared data source for both experimentation and serving, to avoid skew. But even without a feature store, you still have to measure skew, exactly as Zinkevich’s rule #37 says.
The simplest approach is to log feature values at serving time, then recompute them offline for the same day and the same store, and compare the two numbers.
Third, test the infrastructure separately from the ML. Zinkevich writes that infrastructure should be tested independently of the machine learning. For the supermarket chain, you can replace the model with a “dumb” baseline: forecast this week whatever sold last week. Then run the baseline through the whole pipeline, from reading the data and generating forecasts to writing orders into the ordering system.
If the baseline also corrupts the orders, the fault is in the pipeline, and you know it before you wrongly blame the model. Only once the baseline runs cleanly should the real model be swapped in.
Then monitor on real data. Google Cloud defines monitoring as collecting statistics on model performance based on live data. For a forecasting problem, each day when actual sales come in, you compare them with the previous day’s forecast and compute the error by product group. Alongside that, check the inputs too, such as the number of rows received and the share of null values, because broken data usually shows up here first.
Monitoring without alerts is just a dashboard nobody opens. Alert thresholds should be agreed with the warehouse manager, in their language, for example “how far off, in percent, before someone must be told”. A threshold you set yourself in code is one they will not understand and nobody will watch.
Finally, the path back to training. CD4ML stresses that after deployment you must understand how the model behaves in production and feed real data back into the training loop. At the supermarket, daily actual sales are the labels for the next training run. After retraining, there needs to be a review step: the new model replaces the old one only if it performs better on the most recent few weeks of data.
At 2 a.m., which answers do you need ready?
One way to test readiness is to put yourself on the on-call call and ask what each question needs in order to be answered.
| The client’s question | What must already exist |
|---|---|
| Which model is running? | A model version recorded in every log line |
| Is it the data or the model? | Skew measurements and baseline run results |
| When did the model start getting worse? | Performance statistics on real data, with alerts |
| Once fixed, how does it go live? | A retraining pipeline with a review step and rollback |
If any cell in the right-hand column is still empty, that is the work to do before the model is allowed to write a single order.
Common FDE mistakes
One is believing that sharing feature-computation code rules out skew, when the input data sources may already differ. Another is tracking only latency and CPU and assuming that counts as monitoring the model.
The hardest mistake to spot is handing over the system without handing over the alerts. When you leave the site, someone on the client side must know where alerts go and what to do when one arrives.
How to show this skill on a CV
When reading an FDE job description, look for phrases such as monitoring, alerting, upgrades, on-call, production issues. They signal that the company needs someone who can run systems, not just someone who can train models.
On your CV, instead of writing “deployed a forecasting model”, be more specific, for example “built a versioned pipeline, measured skew between training and serving, set alerts on thresholds agreed with the business”.
Before an interview, prepare a true story: the last time your model went wrong on real data, how you found out and how long it took to isolate the fault. A story like that demonstrates exactly the work these job postings describe, more clearly than any accuracy figure.
5 sources
- MLOps: Continuous delivery and automation pipelines in machine learning · 2024-08-28
- Rules of Machine Learning (Martin Zinkevich) · 2025-08-25
- Hidden Technical Debt in Machine Learning Systems · 2015
- Continuous Delivery for Machine Learning (CD4ML) · 2019-09-19
- Forward Deployed Infrastructure Engineer, New Grad - US Government (Palantir)