FDE PulseFDE jobs open 434New in the last 7 days 27
VI

The newspaper of the Forward Deployed Engineer

Guides

Build CI/CD for ML models with CML: post metric comparisons on every pull request

With one workflow file and three Python scripts, reviewers can see whether your change makes the model better or worse before they merge it.

Ảnh kỹ sư phần mềm đang xem code trên laptop hoặc cùng review code trong văn phòng.
Photo: Negative Space / CC0

In brief

  • CI for ML has to check data, schema and the model, not just code.
  • CML runs on GitHub Actions, retrains the model on every PR and posts a metric report as a comment.
  • CML does not ship with DVC. If your data lives in DVC, you need to add the Setup DVC action.
ShareLinkedInFacebookX

In many ML teams, the hardest question in code review has nothing to do with the code. It is whether the change makes the model better or worse. A reviewer sees three edited feature lines in the diff, but nobody knows accuracy has fallen from 0.91 to 0.89 until the model reaches production.

CML, short for Continuous Machine Learning, describes itself as CI/CD for machine learning projects. It promises something specific: every pull request produces its own report with metrics and charts. This guide builds that pipeline on GitHub from start to finish.

For anyone aiming to work as an FDE, this skill pays off right away. At a customer site, don’t open by proposing a large MLOps platform. Start with a small mechanism that runs in the customer’s own repo, so everyone can see how the model changes after each edit.

How does CI for ML differ from CI for web apps?

GitLab defines CI as validating code changes early and often through automated builds and tests. CD automatically prepares tested code so it is always ready to deploy. Google Cloud’s MLOps documentation widens that scope for ML systems: CI must also test the data, the data schema and the model.

Google also separates out continuous training (CT), meaning the automatic retraining and serving of models, and calls it a property unique to ML systems. That is why the pipeline in this guide has three stages: check the data, retrain, then compare metrics. Code that passes its tests can still produce a worse model.

What you will build and what you need

Here is the end result. Whenever someone opens a pull request, GitHub Actions runs a workflow. GitHub’s documentation defines a workflow as an automated process that runs one or more jobs. An event triggers it, meaning a specific activity in the repo, such as opening a pull request.

Your job checks the data, trains the model, compares the results with a baseline and has CML post the results table to the PR.

You need a GitHub repo, Python, a small CSV data file in the repo and a model that trains in a few minutes. The example below is simplified to keep it easy to follow. The column names, libraries and thresholds are placeholders, so replace them with your own.

Step 1: Block bad data before training

Training takes time, so data errors need to be caught first. The script below checks that all required columns are present and that none contain empty values. If either check fails, it exits with a non-zero code and the workflow stops.

# check_data.py (ví dụ giản lược)
import sys
import pandas as pd

EXPECTED = ["tenure", "monthly_spend", "churned"]
df = pd.read_csv("data/train.csv")

missing = [c for c in EXPECTED if c not in df.columns]
if missing:
    sys.exit(f"Thiếu cột: {missing}")
if df[EXPECTED].isnull().any().any():
    sys.exit("Có giá trị rỗng trong cột bắt buộc")
print(f"OK: {len(df)} dòng")

To test it, run python check_data.py locally. Then delete a column from the CSV and run it again. The script should report an error. If you have never seen a check fail, you haven’t tested it.

Step 2: Train and write metrics to a file

The only requirement for train.py is that it writes metrics to a machine-readable file. The training code is up to you. The part to keep is the ending.

# cuối train.py (giản lược)
import json
metrics = {"accuracy": round(acc, 4), "recall": round(rec, 4)}
with open("metrics.json", "w") as f:
    json.dump(metrics, f)

Run it once on the main branch, copy the output to baseline.json and commit it. Every PR is compared against this reference point. It is simpler than comparing against the main branch directly in CI. The cost is that you have to remember to update the baseline when you accept a new model.

Step 3: Generate a Markdown comparison report

CML comments are Markdown, so the comparison script only has to print a table.

# compare.py (giản lược)
import json
base = json.load(open("baseline.json"))
new = json.load(open("metrics.json"))

print("| Metric | Baseline | PR | Chênh lệch |")
print("|---|---|---|---|")
for k in base:
    d = new[k] - base[k]
    flag = "⚠️" if d < -0.01 else ""
    print(f"| {k} | {base[k]} | {new[k]} | {d:+.4f} {flag} |")

Try it with hypothetical numbers: the baseline accuracy is 0.91 and the PR produces 0.89. The difference is -0.02, which crosses the -0.01 threshold, so that row gets a warning flag. The reviewer sees it immediately, without rerunning anyone’s notebook.

Step 4: Wire everything into the workflow

Create the file .github/workflows/cml.yaml. These details come from the CML documentation: use the setup-cml action to install CML, with no Docker container needed. Pass the token through the REPO_TOKEN variable, taken from secrets.GITHUB_TOKEN. The final step calls cml comment create. Everything else is a minimal GitHub Actions skeleton. Check the action version tags against CML’s Get Started page.

name: model-check
on: pull_request
jobs:
  train-and-report:
    runs-on: ubuntu-latest
    steps:
      - uses: actions/checkout@v4
      - uses: iterative/setup-cml@v2   # đối chiếu tag trong tài liệu CML
      - name: Kiểm tra dữ liệu, huấn luyện, báo cáo
        env:
          REPO_TOKEN: ${{ secrets.GITHUB_TOKEN }}
        run: |
          pip install -r requirements.txt
          python check_data.py
          python train.py
          python compare.py > report.md
          cml comment create report.md

The order of commands in run follows the logic of the whole guide: bad data stops the run at the second line, before any time is spent on training.

Step 5: Open a PR and read the comment

Create a branch, change one model parameter, push and open a pull request. The CML documentation describes what happens next in a single line: after a short while, a comment with the CML report appears in the PR. The metric table from Step 3 shows up below the PR description.

Test two more cases. First, a PR that breaks the CSV: the workflow should fail at the data check. Second, a PR that lowers a metric: the workflow should pass, but the comment should carry a warning flag. These are two different kinds of failure, and your team needs to be able to tell them apart.

The most common mistakes

The first is that the data isn’t on the runner. If your data is managed with DVC, note that CML does not ship with DVC or its dependencies. The documentation says you must add a separate Setup DVC action before the data check step.

The second is a workflow that passes but posts no comment. Check that REPO_TOKEN sits in the env of the step that calls cml comment create, then check the token’s permissions in the repo settings.

The third is harder to spot: a stale baseline. If you merge a better model and forget to update baseline.json, every later PR is compared with an outdated reference point and looks better than it really is.

Where this skill comes up at customer sites

According to CML’s homepage, the tool uses GitLab, GitHub or Bitbucket to manage ML experiments and to track who trained a model or changed the data, and when. When a model misbehaves, those are the questions you need to answer: who changed what, and when.

If the PR history can answer them, you don’t have to dig through each person’s notebooks.

One practical tip: in your first week with a customer, hold off on proposing a new MLOps platform. Ask which Git hosting they use, then build exactly one workflow like this for their most important model.

When you read FDE or ML engineer job descriptions and see phrases such as “CI/CD for ML”, “model validation” or “reproducible training”, be ready to walk through a pipeline like this one, with a screenshot of a metric comment on a PR.

The end goal is a customer team that reads the metric table before the diff. Then every pull request answers the hardest review question on its own: is the model better or worse?

7 sources
Read next on the roadmap · Stage 5: DeploymentHands-on: build a Slack approval gate before your agent issues refunds or sends emailsSix steps to make sure a person approves every refund or email your agent sends, without leaving customers waiting forever. Answer the button click within 3 seconds, wait as long as the decision takes, and reject automatically if nobody responds.