# When a model is wrong but reports no errors: build your own data drift detection with Prometheus and Grafana

> A model can keep returning HTTP 200 in a few milliseconds even after customer data has moved far from the data it learned on. This guide shows how to build a system that catches the change before the customer does.

Bản gốc: https://fdetimes.net/en/guides/detect-data-drift-prometheus-grafana/

A fraud-scoring model can return results in a few milliseconds, answer every request with HTTP 200 and write not a single error line to the log. It can still be wrong. When production data has moved significantly away from the training data, which is what data drift means, the model has no way of knowing and nothing to report.

Only a system built specifically to measure the data distribution will notice.

For an FDE, this is part of the daily job. You deploy a model onto a customer's infrastructure, and a few weeks later someone asks: "Why has the model been so poor lately?" If you already have a dashboard showing which feature started to shift and on which day, you answer with data. If you do not, all you can do is guess.

This article builds exactly that system with three tools. Evidently collects and computes the metrics, Prometheus stores them, and Grafana displays them and raises alerts.

## What you will build, and what you need

The end result is two Python processes. The model-serving service exposes metrics on prediction counts and latency through an HTTP endpoint. A drift job that runs on a schedule exposes drift scores through a second endpoint.

Prometheus visits both endpoints at regular intervals to pull the numbers, Grafana has a panel charting each feature's drift score over time, and you receive an alert when a score crosses the threshold.

You need Python 3, a trained model (any scikit-learn classifier will do), and Prometheus and Grafana installed locally or running in containers. You also need a reference dataset, usually the data the model was trained on. Without it there is nothing to compare against.

One architectural point is worth grasping before writing any code. Prometheus collects time series using a pull model over HTTP: each process only has to expose its metrics on an endpoint, and Prometheus comes to fetch them. So you do not write any code to push metrics anywhere.

## Step 1: pick the right metric type for each question

Prometheus provides client libraries for instrumenting application code, and a model-serving service is just another application. To monitor a model you need to answer three questions, and each fits one metric type.

| Question | Metric type | Why |
|---|---|---|
| How many predictions has the model returned? | Counter | Only increases monotonically, which suits counting |
| How long does inference take? | Histogram | Counts each observation into configurable buckets |
| How far is the data drifting? | Gauge | The value can go up or down freely |

The first two questions are answered inside the model-serving service. Below is a minimal sketch using the Python client library. Function names can differ between versions, so check against the docs for the version you have installed.

```python
# serve.py — sketch; check the API against the client library docs
from prometheus_client import Counter, Histogram, start_http_server

PREDICTIONS = Counter("predictions_total", "Number of predictions returned")
LATENCY = Histogram("inference_seconds", "Inference time")

start_http_server(8000)

def predict(x):
with LATENCY.time():
y = model.predict(x)
PREDICTIONS.inc()
return y
```

**Check:** send a few requests to the predict function, then open `localhost:8000/metrics` in a browser. You should see `predictions_total` rise after each request, along with the bucket lines for `inference_seconds`.

## Step 2: compute drift in batches, in a process with its own endpoint

Drift is a property of a distribution, and a single request has no distribution. In the architecture Evidently describes, Evidently reads the model's logs, compares recent data with the reference set, and then exposes an endpoint for Prometheus to collect from.

The "own endpoint" detail matters more than it looks. A Gauge lives in the memory of the process that created it. A scheduled job running in a different process cannot write to the model-serving service's Gauge. So the drift job has to create its own Gauge and expose its own metrics on a different port, here 8001.

You could also run the drift calculation as a background thread inside the service itself, but keeping it separate stops the heavy computation from slowing down inference.

```python
# drift_job.py — pseudocode. compute_drift() stands in for Evidently's Data Drift Report;
# the Evidently API changes between versions, see the current docs.
from prometheus_client import Gauge, start_http_server

DRIFT = Gauge("feature_drift_score", "Drift score per feature", ["feature"])
start_http_server(8001)

while True:
current_window = load_recent_prediction_log()   # read the service's log
for feature in FEATURES:
score = compute_drift(reference[feature], current_window[feature])
DRIFT.labels(feature=feature).set(score)
sleep_until_next_run()                          # e.g. once an hour
```

Picture a fraud model with 20 features. Each feature is one label value on the same Gauge, so you get 20 time series. Evidently notes that this Data Drift example applies in the same way to its other Reports, so later you can add data-quality metrics without changing the architecture.

**Check:** after the first run, `localhost:8001/metrics` should contain 20 lines of `feature_drift_score{feature="..."}`, each carrying a value.

## Step 3: configure Prometheus to pull the metrics

In the configuration file, `scrape_interval` sets how often Prometheus pulls metrics. Because there are two processes, you declare two targets. The file below is trimmed for illustration.

```yaml
# prometheus.yml (trimmed)
global:
scrape_interval: 15s
scrape_configs:
- job_name: "fraud-model"
static_configs:
- targets: ["localhost:8000"]
- job_name: "drift-monitor"
static_configs:
- targets: ["localhost:8001"]
```

A quick calculation. At a 15-second interval, each time series receives 240 samples an hour. But if the drift job runs only once an hour, 239 of those samples repeat the same value. That is not wrong, but you need to understand it when reading the chart: the drift line moves in steps that follow the job's schedule, not the scrape interval.

Start Prometheus with this file (the Getting Started page has the right command for your version), open the web interface and type the simplest possible PromQL query, the metric name `feature_drift_score`.

**Check:** the targets page should show both `fraud-model` and `drift-monitor` as up. If a job is missing, the cause is usually a wrong port or a process that is not running.

## Step 4: build the panel and alerts in Grafana

In Grafana, add Prometheus as a data source, create a time series panel and use the query `feature_drift_score`. Each feature becomes its own line. This is the chart you will open when the customer asks "is something wrong with the model?"

Next come alerts. Grafana lets you set alerts via email, Slack or SMS on custom thresholds. DataCamp's tutorial uses a threshold of 0.026, meaning the system sends an alert when the drift score exceeds that level. To try it out, you can write a condition such as `feature_drift_score > 0.026`.

**Check:** take the test data, double the values of one feature and feed it into the current data window. That feature's line should jump and the alert should arrive on the channel you configured.

**Điểm mấu chốt:** The 0.026 figure is an example for learning how to build an alert, not the right threshold for a customer's data.

## The right threshold comes from the customer's own data

The safe approach is to run the system in observe-only mode for a few weeks while the model is performing well, then set the threshold from those numbers. Suppose the drift job runs every hour for 4 weeks: each feature gets 4 × 7 × 24 = 672 baseline points.

Sort those 672 points and take the 99th percentile. Since 1% of 672 is 6.72, this value sits around the seventh-highest point. Say it is 0.04: that means in 99% of normal hours the drift score did not exceed 0.04.

You can set the threshold somewhat higher, say 0.05, so that only genuinely abnormal shifts trigger an alert. The numbers here are hypothetical, but the method works on real data.

Calculate it separately for each important feature, because some features naturally fluctuate more than others. In PromQL you can use `quantile_over_time` over the baseline period. Check the syntax in the Prometheus docs before using it.

## Mistakes that make a monitoring system useless

The first is using a Counter for drift scores. A Counter only goes up, so when drift falls the metric cannot reflect it and your chart will mislead you. Any value that represents "current state", such as a drift score or accuracy, must use a Gauge.

The second is copying a threshold straight from a tutorial. A threshold that is too low floods the customer's Slack with alerts, and after a week nobody reads them. The percentile calculation above costs a few lines of code and saves you from this.

The third is labelling by something with a very large number of values, such as customer ID. As in step 2, each label value produces its own time series: 20 features means 20 readable lines, while a per-user label makes the number of time series grow with the number of users.

Keep labels to dimensions with a small, known set of values, and read the Prometheus docs carefully before adding a new label.

The last one few people notice: the reference set no longer matches the model in production, because the model has been retrained while the reference set was left unchanged.

## Where this skill shows up in customer work

In customer work, the hardest part is usually not the code. You need to sit down with the business side and agree on three things: which features matter enough to warrant their own alert, who receives the alerts, and what happens next when an alert fires. Without a follow-up step, the dashboard is decoration.

So in your first session with a customer, ask whether they use Prometheus, Grafana or some other monitoring system. Plugging the model's metrics into infrastructure they already have is usually far easier to get accepted than proposing a new stack.

When reading FDE or MLOps job descriptions, look for phrases such as "model monitoring", "observability" or "drift detection". On a CV, "set up Prometheus and Grafana" says little. "Detected drift across 20 features, alerting via Slack, with thresholds taken from the 99th percentile of baseline data" tells a recruiter you understand the problem.

Every deployed model will go wrong at some point. The difference is whether you find out from an alert as the data starts to shift, or wait until the customer calls to tell you.

**Thử ngay tuần này:**

- Take an existing scikit-learn model, wrap it in a service with a /metrics endpoint carrying a Counter for predictions and a Histogram for latency, then write a separate drift job with a second endpoint holding a Gauge of per-feature drift scores.
- Take the test data, deliberately shift the distribution of one feature (for example, double the amount), and see whether the Grafana panel and the alert fire.
- Write a short README explaining which percentile of the baseline data you used for the alert threshold, and add the repo link to your CV under MLOps projects.

## Nguồn

- [Prometheus Overview](https://prometheus.io/docs/introduction/overview/)

- [Getting Started with Prometheus](https://prometheus.io/docs/tutorials/getting_started/)

- [Metric types | Prometheus](https://prometheus.io/docs/concepts/metric_types/)

- [Evidently and Grafana: ML monitoring live dashboards](https://evidentlyai.com/blog/evidently-and-grafana-ml-monitoring-live-dashboards)

- [Grafana Tutorial: Monitoring Machine Learning Models (DataCamp)](https://www.datacamp.com/tutorial/grafana-tutorial-monitoring-machine-learning-models)
