# Hands-on: five layers of tests that stop dirty client data before it reaches the model

> On a client site, the problem is rarely a bug in the model. More often it is a CSV file that writes amounts a different way on each row and leaves some rows blank. This guide builds the testing layers one at a time, so a file like that is stopped at the start of the pipeline.

Original: https://fdetimes.net/en/guides/data-pipeline-testing-five-layers/

Picture your first week at a retail chain. The client sends an orders file exported from three systems. In the `amount` column, some rows read "1.250.000", some read "1,250,000", some are blank, and one order ID appears twice. The revenue forecasting model will accept every one of those rows without complaint.

IBM sums up the problem with the familiar phrase "garbage in, garbage out": if the input data is poor, the model's output will be wrong too. IBM also defines data quality as the degree to which a dataset meets the criteria of **accuracy, completeness, validity and consistency**. This guide turns those four criteria into code you can run.

For a forward deployed engineer (FDE), this work often decides whether the first demo keeps the client's trust. When the model returns a strange number, you need to show which row caused it and why that row got through.

## What will you build, and what do you need?

You will build five layers of checks based on the testing pyramid. The bottom layer holds many fast, cheap unit tests. The top layer holds very few end-to-end (E2E) tests, because they are slow, expensive and fragile. BrowserStack stresses that the pyramid suggests where tests belong; it does not require a fixed ratio.

All you need is Python 3 and an empty folder. The code below deliberately uses plain Python, with `assert` and the `csv` module, so the logic stays visible. This is a **simplified** version: the pipeline only accepts orders in Vietnamese dong (VND).

The first five steps match the five layers. Step 6 is not a sixth layer. It is a change of tooling: the data checks and the gate move to a dedicated library as the project grows.

## Step 1: unit tests for pure transform functions

Start with the function that normalises amounts, which is where errors most easily get in. The function below is **for VND only**. It strips every dot and comma, because the dong has no decimal fraction.

```python
# transform.py
def normalize_amount(raw):
    """VND only. Do not use for currencies with a decimal part."""
    if raw is None:
        return None
    cleaned = raw.strip().replace(".", "").replace(",", "")
    if not cleaned.isdigit():
        return None
    return float(cleaned)
```

```python
# test_transform.py
from transform import normalize_amount

assert normalize_amount("1.250.000") == 1250000.0
assert normalize_amount("1,250,000") == 1250000.0
assert normalize_amount(" 500000 ") == 500000.0
assert normalize_amount("") is None
assert normalize_amount("abc") is None
print("unit OK")
```

Run `python test_transform.py`. If you see `unit OK`, the tests pass. Guru99 explains why this layer matters: unit tests expose errors close to where they start, so fixes are faster and cheaper.

The VND-only limit is deliberate. Pass a USD order of "99.50" to this function and the result is 9950, wrong by a factor of 100, with no error raised. That is why, in step 2, every non-VND row is flagged and quarantined until you write a separate normalisation function for each currency.

Also try `normalize_amount("-50000")`. The result is `None`, because the minus sign makes `isdigit()` return false. Whether refund orders may be negative is a business decision. Ask the client rather than guess.

## Step 2: turn the four quality criteria into checks

Unit tests check your code. This step checks the client's data. Each condition below maps to one of IBM's criteria. The error messages are in Vietnamese: "thiếu order_id" means missing order_id, "order_id trùng" duplicate order_id, "amount không hợp lệ" invalid amount and "currency chưa hỗ trợ" unsupported currency.

```python
# checks.py
SUPPORTED_CURRENCY = "VND"  # simplified version: VND only

def check_batch(rows):
    failures, seen = [], set()
    for i, r in enumerate(rows):
        oid = r.get("order_id")
        if not oid:
            failures.append((i, "missing order_id"))          # completeness
        elif oid in seen:
            failures.append((i, "duplicate order_id"))        # consistency
        seen.add(oid)
        if r.get("amount") is None or r["amount"] <= 0:
            failures.append((i, "invalid amount"))            # validity
        if r.get("currency") != SUPPORTED_CURRENCY:
            failures.append((i, "unsupported currency"))      # validity
    return failures
```

(The comments mark the criteria: đầy đủ is completeness, nhất quán is consistency and hợp lệ is validity.)

Accuracy is the hardest criterion to write as a rule, because a number can be well formed and still not match reality. In practice, ask the client for a reference figure, such as last month's total revenue in the accounting system. Then add a check that compares the batch total with that figure.

## Step 3: integration tests between stages

Functions that each work correctly can still fail when combined. Guru99 describes integration testing as the layer that finds errors where modules interact. In a pipeline, that means the joins between the read, transform and check steps.

Create `fixtures/orders_sample.csv` with five rows and two planted errors:

```text
order_id,amount,currency
A1,1.250.000,VND
A2,"1,250,000",VND
A2,300000,VND
A4,,VND
A5,99000,VND
```

```python
# test_pipeline.py
import csv
from transform import normalize_amount
from checks import check_batch

def load(path):
    with open(path, newline="", encoding="utf-8") as f:
        rows = list(csv.DictReader(f))
    for r in rows:
        r["amount"] = normalize_amount(r.get("amount"))
    return rows

rows = load("fixtures/orders_sample.csv")
msgs = {m for _, m in check_batch(rows)}
assert len(rows) == 5
assert msgs == {"duplicate order_id", "invalid amount"}
print("integration OK")
```

This test catches a type of error that unit tests miss. In the file, an amount containing commas must be wrapped in double quotes. If a later export drops the quotes, the `csv` module splits "1,250,000" across several columns, the `currency` column receives an unexpected value, and the test fails at once instead of letting the wrong number through.

## Step 4: a gate for every run

According to Guru99, a smoke test confirms that a build is stable enough for further testing. In a pipeline, apply the same idea to each batch: a batch that is clean enough passes, and anything else stops before it reaches the model.

```python
# gate.py
from checks import check_batch

MAX_BAD_RATIO = 0.02  # assumed threshold, agree on it with the client

def gate(rows):
    bad = {i for i, _ in check_batch(rows)}
    ratio = len(bad) / max(len(rows), 1)
    if not rows or ratio > MAX_BAD_RATIO:
        raise SystemExit(f"BLOCKED batch: {len(bad)}/{len(rows)} bad rows")
    clean = [r for i, r in enumerate(rows) if i not in bad]
    quarantine = [r for i, r in enumerate(rows) if i in bad]
    return clean, quarantine
```

(The 2% threshold is an assumption to agree with the client. "CHẶN batch: … dòng lỗi" means "batch blocked: … bad rows".)

Suppose a batch has 1,000 rows. If 15 rows are bad, the error rate is 1.5%: the batch goes ahead with 985 clean rows, and the 15 bad rows go to quarantine for the client to review. If 37 rows are bad, the rate is 3.7% and the whole batch is blocked, because that many errors usually means the source system has changed, not just that someone made a few typos.

**Key point:** Stopping a dirty batch at the door costs far less than explaining why the model got it wrong.

## Step 5: a few E2E tests, run close to production

The Microsoft Code-With Engineering Playbook describes E2E testing as checking the whole application from start to finish with every component connected. The playbook also warns that the higher a test sits in the pyramid, the slower it runs and the more effort it takes to write, run and maintain.

One or two scenarios are therefore enough: take a real file with sensitive data masked, run it through the whole pipeline, and check whether the model's output makes sense.

TestGrid advises always putting the user's perspective first. On a client site, that means running E2E tests as close to production as possible, with the same kind of connection, the same access permissions and the same file format the real system will send. A sample file you wrote on your laptop cannot stand in for those.

## Step 6: move to Great Expectations

This step adds no new layer. It replaces the tooling for the data checks (step 2) and the gate (step 4). Once the number of rules grows, move the data checks to Great Expectations.

GX Core is a Python library for building and running data validation workflows in code. Each condition in `check_batch` maps to an **Expectation**, a verifiable assertion about the data.

The snippet below is an **illustration** of a single rule, "`order_id` must not be missing", written in the GX Core 1.x style. Function names may differ between versions, so check the GX documentation for the version you have installed before you use it.

```python
# gx_demo.py  (illustrative; check against the GX docs)
import great_expectations as gx
import pandas as pd

df = pd.read_csv("fixtures/orders_sample.csv", dtype=str)

context = gx.get_context()
source = context.data_sources.add_pandas("orders")
asset = source.add_dataframe_asset(name="orders_df")
batch_def = asset.add_batch_definition_whole_dataframe("whole")
batch = batch_def.get_batch(batch_parameters={"dataframe": df})

exp = gx.expectations.ExpectColumnValuesToNotBeNull(column="order_id")
print(batch.validate(exp).success)
```

With the fixture from step 3, the result should be `True`, because all five rows have an `order_id`. Delete the ID from one row and run it again to see the result change to `False`.

The `gate` function maps to a **Checkpoint**. The GX documentation describes a Checkpoint as the primary way to validate data when GX is deployed in production. Your next exercise: rewrite the other three rules from step 2 as Expectations, then group them into a Checkpoint.

## Common mistakes

The most common mistake is putting all the effort into E2E tests and skipping unit tests. Then every failure costs you an afternoon spent working out which step broke. Another is setting the blocking threshold without asking the client. A 2% threshold is a business question, not a technical one.

The last is quietly deleting bad rows. Quarantine means keeping the bad rows, recording the reason and sending the list to the client. The client may well find that their own system is producing the bad data.

## How to put this skill on your CV

Do not just write "wrote tests" on your CV. Write a specific line, for example: built a batch validation gate that quarantines bad rows and blocks batches over the error threshold before they reach the model.

At interviews, bring the five-row fixture file from step 3 and be ready to run it on your laptop. Show the interviewer which errors you planted in the data and which layer of the pipeline catches each one.

**Try this week:**

- Pick a data-cleaning function in your current project and write five asserts for it, with at least two bad inputs
- Create a five-row fixture file with two planted errors, then write an integration test asserting that the pipeline catches exactly those two
- Write a gate function with an error-rate threshold, run it on a real batch, and record how many rows were quarantined

## Sources

- [What is Data Quality? (IBM)](https://www.ibm.com/think/topics/data-quality)

- [Testing Pyramid for Test Automation (BrowserStack)](https://www.browserstack.com/guide/testing-pyramid-for-test-automation)

- [Unit Testing Tutorial (Guru99)](https://www.guru99.com/unit-testing-guide.html)

- [Integration Testing Tutorial (Guru99)](https://www.guru99.com/integration-testing.html)

- [Smoke Testing (Guru99)](https://www.guru99.com/smoke-testing.html)

- [End-to-End Testing (Code-With Engineering Playbook)](https://microsoft.github.io/code-with-engineering-playbook/automated-testing/e2e-testing/)

- [End to End Testing: Importance, Process, Best Practices & Frameworks (TestGrid)](https://testgrid.io/blog/end-to-end-testing-a-detailed-guide/)

- [GX Core overview (Great Expectations docs)](https://docs.greatexpectations.io/docs/core/introduction/gx_overview)
