FDE PulseFDE jobs open 441New in 7 days 29Companies hiring 47Remote-friendly 24%Median US pay $216kTop hirer Databricks 125
VI

The newspaper of the Forward Deployed Engineer

Guides

Hands-on: five layers of tests that stop dirty client data before it reaches the model

On a client site, the problem is rarely a bug in the model. More often it is a CSV file that writes amounts a different way on each row and leaves some rows blank. This guide builds the testing layers one at a time, so a file like that is stopped at the start of the pipeline.

In brief

  • Write many cheap unit tests for transform functions and keep only a few expensive E2E tests, because E2E tests break easily.
  • Check data for accuracy, completeness, validity and consistency, then combine those checks into a gate for each batch.
  • In Great Expectations, each assertion about the data is an Expectation, and a Checkpoint is how checks run in production.
ShareLinkedInFacebookX
GraphicFive test layers that stop dirty data
  1. Near-production E2EA few scenarios run the full pipeline; this layer is slow, costly and brittle
  2. Per-batch gateBatches over the error threshold stop; bad rows go to quarantine
  3. Integration between stagesFixture files with planted errors run through read, transform and check steps
  4. Data validationFour criteria: accurate, complete, valid, consistent
  5. Unit tests for transformsMany fast, cheap tests; catch errors close to where they arise

Write many cheap tests at the lower layers and keep only a few costly E2E tests at the top.

Graphic: FDE Times

Picture your first week at a retail chain. The client sends an orders file exported from three systems. In the amount column, some rows read “1.250.000”, some read “1,250,000”, some are blank, and one order ID appears twice. The revenue forecasting model will accept every one of those rows without complaint.

IBM sums up the problem with the familiar phrase “garbage in, garbage out”: if the input data is poor, the model’s output will be wrong too. IBM also defines data quality as the degree to which a dataset meets the criteria of accuracy, completeness, validity and consistency. This guide turns those four criteria into code you can run.

For a forward deployed engineer (FDE), this work often decides whether the first demo keeps the client’s trust. When the model returns a strange number, you need to show which row caused it and why that row got through.

What will you build, and what do you need?

You will build five layers of checks based on the testing pyramid. The bottom layer holds many fast, cheap unit tests. The top layer holds very few end-to-end (E2E) tests, because they are slow, expensive and fragile. BrowserStack stresses that the pyramid suggests where tests belong; it does not require a fixed ratio.

All you need is Python 3 and an empty folder. The code below deliberately uses plain Python, with assert and the csv module, so the logic stays visible. This is a simplified version: the pipeline only accepts orders in Vietnamese dong (VND).

The first five steps match the five layers. Step 6 is not a sixth layer. It is a change of tooling: the data checks and the gate move to a dedicated library as the project grows.

Step 1: unit tests for pure transform functions

Start with the function that normalises amounts, which is where errors most easily get in. The function below is for VND only. It strips every dot and comma, because the dong has no decimal fraction.

# transform.py
def normalize_amount(raw):
    """VND only. Do not use for currencies with a decimal part."""
    if raw is None:
        return None
    cleaned = raw.strip().replace(".", "").replace(",", "")
    if not cleaned.isdigit():
        return None
    return float(cleaned)
# test_transform.py
from transform import normalize_amount

assert normalize_amount("1.250.000") == 1250000.0
assert normalize_amount("1,250,000") == 1250000.0
assert normalize_amount(" 500000 ") == 500000.0
assert normalize_amount("") is None
assert normalize_amount("abc") is None
print("unit OK")

Run python test_transform.py. If you see unit OK, the tests pass. Guru99 explains why this layer matters: unit tests expose errors close to where they start, so fixes are faster and cheaper.

The VND-only limit is deliberate. Pass a USD order of “99.50” to this function and the result is 9950, wrong by a factor of 100, with no error raised. That is why, in step 2, every non-VND row is flagged and quarantined until you write a separate normalisation function for each currency.

Also try normalize_amount("-50000"). The result is None, because the minus sign makes isdigit() return false. Whether refund orders may be negative is a business decision. Ask the client rather than guess.

Step 2: turn the four quality criteria into checks

Unit tests check your code. This step checks the client’s data. Each condition below maps to one of IBM’s criteria. The error messages are in Vietnamese: “thiếu order_id” means missing order_id, “order_id trùng” duplicate order_id, “amount không hợp lệ” invalid amount and “currency chưa hỗ trợ” unsupported currency.

# checks.py
SUPPORTED_CURRENCY = "VND"  # simplified version: VND only

def check_batch(rows):
    failures, seen = [], set()
    for i, r in enumerate(rows):
        oid = r.get("order_id")
        if not oid:
            failures.append((i, "missing order_id"))          # completeness
        elif oid in seen:
            failures.append((i, "duplicate order_id"))        # consistency
        seen.add(oid)
        if r.get("amount") is None or r["amount"] <= 0:
            failures.append((i, "invalid amount"))            # validity
        if r.get("currency") != SUPPORTED_CURRENCY:
            failures.append((i, "unsupported currency"))      # validity
    return failures

(The comments mark the criteria: đầy đủ is completeness, nhất quán is consistency and hợp lệ is validity.)

Accuracy is the hardest criterion to write as a rule, because a number can be well formed and still not match reality. In practice, ask the client for a reference figure, such as last month’s total revenue in the accounting system. Then add a check that compares the batch total with that figure.

Step 3: integration tests between stages

Functions that each work correctly can still fail when combined. Guru99 describes integration testing as the layer that finds errors where modules interact. In a pipeline, that means the joins between the read, transform and check steps.

Create fixtures/orders_sample.csv with five rows and two planted errors:

order_id,amount,currency
A1,1.250.000,VND
A2,"1,250,000",VND
A2,300000,VND
A4,,VND
A5,99000,VND
# test_pipeline.py
import csv
from transform import normalize_amount
from checks import check_batch

def load(path):
    with open(path, newline="", encoding="utf-8") as f:
        rows = list(csv.DictReader(f))
    for r in rows:
        r["amount"] = normalize_amount(r.get("amount"))
    return rows

rows = load("fixtures/orders_sample.csv")
msgs = {m for _, m in check_batch(rows)}
assert len(rows) == 5
assert msgs == {"duplicate order_id", "invalid amount"}
print("integration OK")

This test catches a type of error that unit tests miss. In the file, an amount containing commas must be wrapped in double quotes. If a later export drops the quotes, the csv module splits “1,250,000” across several columns, the currency column receives an unexpected value, and the test fails at once instead of letting the wrong number through.

Step 4: a gate for every run

According to Guru99, a smoke test confirms that a build is stable enough for further testing. In a pipeline, apply the same idea to each batch: a batch that is clean enough passes, and anything else stops before it reaches the model.

# gate.py
from checks import check_batch

MAX_BAD_RATIO = 0.02  # assumed threshold, agree on it with the client

def gate(rows):
    bad = {i for i, _ in check_batch(rows)}
    ratio = len(bad) / max(len(rows), 1)
    if not rows or ratio > MAX_BAD_RATIO:
        raise SystemExit(f"BLOCKED batch: {len(bad)}/{len(rows)} bad rows")
    clean = [r for i, r in enumerate(rows) if i not in bad]
    quarantine = [r for i, r in enumerate(rows) if i in bad]
    return clean, quarantine

(The 2% threshold is an assumption to agree with the client. “CHẶN batch: … dòng lỗi” means “batch blocked: … bad rows”.)

Suppose a batch has 1,000 rows. If 15 rows are bad, the error rate is 1.5%: the batch goes ahead with 985 clean rows, and the 15 bad rows go to quarantine for the client to review. If 37 rows are bad, the rate is 3.7% and the whole batch is blocked, because that many errors usually means the source system has changed, not just that someone made a few typos.

Step 5: a few E2E tests, run close to production

The Microsoft Code-With Engineering Playbook describes E2E testing as checking the whole application from start to finish with every component connected. The playbook also warns that the higher a test sits in the pyramid, the slower it runs and the more effort it takes to write, run and maintain.

One or two scenarios are therefore enough: take a real file with sensitive data masked, run it through the whole pipeline, and check whether the model’s output makes sense.

TestGrid advises always putting the user’s perspective first. On a client site, that means running E2E tests as close to production as possible, with the same kind of connection, the same access permissions and the same file format the real system will send. A sample file you wrote on your laptop cannot stand in for those.

Step 6: move to Great Expectations

This step adds no new layer. It replaces the tooling for the data checks (step 2) and the gate (step 4). Once the number of rules grows, move the data checks to Great Expectations.

GX Core is a Python library for building and running data validation workflows in code. Each condition in check_batch maps to an Expectation, a verifiable assertion about the data.

The snippet below is an illustration of a single rule, “order_id must not be missing”, written in the GX Core 1.x style. Function names may differ between versions, so check the GX documentation for the version you have installed before you use it.

# gx_demo.py  (illustrative; check against the GX docs)
import great_expectations as gx
import pandas as pd

df = pd.read_csv("fixtures/orders_sample.csv", dtype=str)

context = gx.get_context()
source = context.data_sources.add_pandas("orders")
asset = source.add_dataframe_asset(name="orders_df")
batch_def = asset.add_batch_definition_whole_dataframe("whole")
batch = batch_def.get_batch(batch_parameters={"dataframe": df})

exp = gx.expectations.ExpectColumnValuesToNotBeNull(column="order_id")
print(batch.validate(exp).success)

With the fixture from step 3, the result should be True, because all five rows have an order_id. Delete the ID from one row and run it again to see the result change to False.

The gate function maps to a Checkpoint. The GX documentation describes a Checkpoint as the primary way to validate data when GX is deployed in production. Your next exercise: rewrite the other three rules from step 2 as Expectations, then group them into a Checkpoint.

Common mistakes

The most common mistake is putting all the effort into E2E tests and skipping unit tests. Then every failure costs you an afternoon spent working out which step broke. Another is setting the blocking threshold without asking the client. A 2% threshold is a business question, not a technical one.

The last is quietly deleting bad rows. Quarantine means keeping the bad rows, recording the reason and sending the list to the client. The client may well find that their own system is producing the bad data.

How to put this skill on your CV

Do not just write “wrote tests” on your CV. Write a specific line, for example: built a batch validation gate that quarantines bad rows and blocks batches over the error threshold before they reach the model.

At interviews, bring the five-row fixture file from step 3 and be ready to run it on your laptop. Show the interviewer which errors you planted in the data and which layer of the pipeline catches each one.

Was this article useful?

Use with your AI assistantAsk Claude ↗Ask ChatGPT ↗
8 sources
Read next on the roadmap · Stage 5: DeploymentHands-on: an incremental pipeline you can rerun without duplicating dataClients will give you three years of old data and a nightly job. Both are only safe if the second run gives exactly the same result as the first.