# Data lineage: tracing a wrong prediction back, step by step

> A client sends a screenshot of a wrong number. If you trace that number back to the job and the run that produced it, you can look for the root cause methodically instead of guessing at the model.

Bản gốc: https://fdetimes.net/en/guides/data-lineage-debug-wrong-predictions/

On Monday a client posts a screenshot in the shared channel with one line of text: "This week's forecast for store 12 came in below actual sales, and the warehouse team under-ordered." The model has not changed and nobody has redeployed any code, yet the number is wrong. You are the FDE on the project, and the whole room is waiting for you to say where the problem is.

Most people suspect the model first. They retrain, tune hyperparameters and compare metrics again. When neither the model nor the code has changed, as in this case, it makes more sense to suspect something that changed upstream of the model. Data lineage is the skill that lets you find that change systematically.

## Lineage is a map for working backwards

IBM defines data lineage as the process of tracking the flow of data over time: where it originated, how it changed and where it went. One of its main uses is tracing errors back to their root cause, and that level of detail is extremely useful when you are debugging data problems.

The store 12 complaint is exactly the kind of problem lineage exists to solve.

To work backwards you need a simple mental model. OpenLineage, an open framework and specification for collecting lineage, describes every pipeline with three common entities: dataset, job and run. A dataset is a table or a file. A job is a piece of processing that reads one dataset and writes another.

A run is one specific execution of a job. It happened at a particular time, and its input data may differ from the previous run's.

Think in these three entities even when the client's pipeline has no lineage tooling at all. Every step backwards answers three questions: which dataset holds the number, which job wrote it and which run produced it.

## Tracing store 12's number

What follows is a hypothetical example, written in enough detail for you to follow along. The client's system has four steps. Point-of-sale (POS) data is loaded into the `sales_daily` table. A job computes the feature `avg_sales_7d`, the average quantity sold over seven days. The forecasting service reads that feature at serving time, and a dashboard displays the `forecast_qty` column.

First, pin down the wrong number exactly, because "the forecast is low" is too vague. You need to know the row, meaning store 12, the SKU and the week, and which run of the forecasting service produced it. Until you have the run id, every comparison you make afterwards is guesswork.

Then work backwards one step at a time. At each step, compare the actual value with the value you expect:

| Step (working backwards) | Question | Result in the example |
|---|---|---|
| Dashboard `forecast_qty` | Does it match the service output? | Yes, so the dashboard is not at fault |
| Forecasting service, Monday's run | What was the input feature? | `avg_sales_7d` = 90 |
| Feature job | Recomputing by hand from `sales_daily`, what do you get? | Still 90, so the formula is correct |
| `sales_daily` table, `qty` column | What does this column mean, and since when? | Since last week, `qty` has had returns subtracted |

Now the cause is clear. Say the store sells 100 units a day and 10 units are returned. When the model was trained, `qty` was gross sales, so the feature hovered around 100. The POS team then changed `qty` to net sales, so at serving time the feature is only 90.

The model is still faithful to what it learned. Its input now means something different.

**Điểm mấu chốt:** The model is not broken. An upstream column changed meaning, and the model still assumes the old meaning.

## Why this kind of bug is hard to see

Snowflake puts training-serving skew among the pipeline bugs that are hard to diagnose, and the example shows why. No job failed and no table was empty. Every step did exactly what its code says. The only thing that changed was a definition.

Table-level lineage only tells you that `sales_daily` sits upstream of the model, and that is not enough. You need column-level lineage, which follows dependencies through the whole pipeline and links an upstream change to the downstream features and training data it affects.

Only at column level can you see the path from `qty` to `avg_sales_7d` to `forecast_qty`.

The long-term fix is a single standard feature definition used for both training and serving. A feature store helps most when several models share features or when inference needs low latency.

If the client runs a single weekly batch model, a shared feature-definition module may be all you need.

## If you fix the source, who else is affected?

This is where newer FDEs tend to slip. You have found the cause, and you want to ask the POS team to change `qty` back to its old meaning. But the finance team may depend on that net figure for its reports.

Lineage has a second use: showing what a specific change will affect, known as impact analysis. Before you propose a fix, work forwards from the `qty` column and list every model, dashboard and table that reads it. Usually the sensible option is to add a new column such as `qty_gross` and point the feature at it, rather than changing the column's meaning yet again.

## Do it yourself: five steps for the next incident

1. **Pin down the coordinates.** Turn the complaint into an exact location: which row, which time, which run.
2. **Work backwards one step at a time.** At each step, record the dataset, the job and the run.
3. **Compare against expectations.** Recompute the expected value at each step and compare it with the actual value. The first step that does not match is where you dig.
4. **Compare the two feature definitions.** Put the training-time definition next to the serving-time definition and compare them line by line, because skew usually hides in the differences between them.
5. **Work forwards before you fix anything.** Run an impact analysis before you change anything, and once it is fixed, leave lineage in place for next time.

On tooling: Airflow can track lineage between tasks and send hook-level lineage to a central collector. However, Airflow's own documentation says the feature is highly experimental and subject to change.

So do not assume the graph Airflow produces is complete. If the client uses a mix of tools, OpenLineage is a vendor-neutral standard worth considering.

## Common traps

The most common trap is retraining before tracing back. If the input has changed meaning, the new model simply learns the new meaning and hides the real cause. A related habit is stopping at table-level lineage, where you know which tables are involved but not which column changed.

If you forget to record the run id, you end up comparing today's data with a prediction produced from yesterday's. And if you fix the source without working forwards, you resolve this incident and cause another one for the team next door.

An incident like store 12 also makes some of the best material for a CV if you are trying to move into an FDE role. Rather than writing "experienced in debugging pipelines", describe how you traced four steps back from the dashboard to the `qty` column, found that it had switched from gross to net, and checked downstream before proposing a new `qty_gross` column.

The next time a client sends a screenshot, open the lineage map, not a notebook for retraining. The client needs to know where the wrong number came from, and retraining usually will not tell them.

**Thử ngay tuần này:**

- Pick a number on a dashboard, or a prediction your team serves, and sketch its column-level lineage by hand. For each step, write down the job name and how to get the run id.
- For each feature of one model, compare the training-side definition with the serving-side definition. Note every place where the two sides compute it with different code.
- Read the dataset/job/run section of the OpenLineage documentation and try mapping one of your existing Airflow pipelines onto those three concepts.

## Nguồn

- [What Is Data Lineage? | IBM](https://www.ibm.com/think/topics/data-lineage)

- [AI Data Pipelines: Why Data Consistency Matters as Much as the Model (Snowflake)](https://www.snowflake.com/guides/what-feature-store-machine-learning/)

- [Lineage — Airflow 3.3.2 Documentation](https://airflow.apache.org/docs/apache-airflow/stable/administration-and-deployment/lineage.html)

- [About OpenLineage | OpenLineage](https://openlineage.io/docs/)
