FDE PulseFDE jobs open 434New in the last 7 days 27
VI

The newspaper of the Forward Deployed Engineer

Guides

DVC in practice: make every new client data drop a traceable commit

When a client asks "what data was last month's model trained on?", an FDE should answer with a commit hash, not from memory.

Kỹ sư ngồi trước laptop hiển thị mã nguồn và dữ liệu, không khí làm việc tập trung tại văn phòng.
Photo: Negative Space / CC0

In brief

  • Git holds a small .dvc file containing an MD5 hash, the size and the file count. The actual data lives on a remote such as S3.
  • Every time you run git checkout, also run dvc checkout so the data matches the code.
  • dvc repro skips unchanged stages. Committing dvc.lock to Git pins down a reproducible run.
ShareLinkedInFacebookX
GraphicHow a new data drop moves through DVC and Git
  1. 1dvc addCreates a small .dvc file with the data's MD5 hash, size and file count
  2. 2git commit the .dvc fileData metadata is versioned alongside the source code
  3. 3dvc pushUploads the actual data to a remote the client approves: S3, Azure, GCS, SSH
  4. 4dvc reproReruns only stages whose inputs changed; commit dvc.lock to pin the run
  5. 5git checkout + dvc checkoutRestores the exact data and code behind an old model

Each client data drop becomes a commit. The data sits on a remote and can be restored at any time.

Graphic: FDE Times

On Monday the client sends this week’s data file. On Wednesday they send a “fixed” version. By Friday the new model is performing worse, and the first question in the meeting is: which version was the production model trained on?

If your answer is “probably Wednesday’s”, you have just lost some of the client’s trust. FDEs run into this all the time when deploying ML in the field. Data keeps changing, and while the code is in Git, the data usually sits in a folder as something like final_v3_moi.csv (“final v3 new”).

This guide walks through a workflow you can set up on a laptop in one evening using DVC. The DVC homepage sums up its philosophy as managing data the same way people manage code, and the Get Started docs call it “Git for data” outright.

What will you build, and what do you need first?

Picture a retail client that sends a data/data.xml file every week. The exercise has three goals: each data drop maps to a Git commit, the actual data lives in storage the client has approved, and you can get back any drop whenever you need it.

You need Git, DVC installed, and a Git repo initialised for DVC as described on the Get Started page. One point to understand before typing any commands: technically, DVC itself is not a version control system. It leaves the history to Git, and the contents of the .dvc file determine which version of the data is in effect.

Steps 1–2: turn the first data drop into a commit

Add the file to DVC:

dvc add data/data.xml

DVC creates a metadata file, data/data.xml.dvc. Open it and take a look. It records the size, the number of files and, most importantly, an MD5 hash. That hash identifies the data version.

Then commit the metadata file to Git, not the data itself:

git add data/data.xml.dvc
git commit -m "Dữ liệu khách: đợt tuần 1"

(The commit message reads “Client data: week 1 batch”.) From this point on, the data’s metadata sits in Git history right next to the source code. The example above is trimmed down. When you run dvc add, read DVC’s output, because it may suggest other files to git add.

Step 3: put the data where the client allows

If Git holds only a small file, where does the actual data go? DVC uses remotes: external storage for data and models, with support for S3, Azure, GCS, SSH and many others. This matters to FDEs because the client usually decides where its data is allowed to live.

dvc remote add myremote s3://mybucket
dvc push

dvc push uploads the data in the cache to the configured remote. This is a simplified example. Before you configure a real remote for a client, read the Remote Storage page in the DVC docs carefully.

To check that it worked, open the bucket. The data you just pushed should be there. If the client only allows an internal server over SSH, you change the remote URL, not the workflow.

Step 4: the second data drop arrives

The client sends a new version. You overwrite the file and repeat the same routine:

dvc add data/data.xml
git add data/data.xml.dvc
git commit -m "Dữ liệu khách: đợt tuần 2"
dvc push

Open data/data.xml.dvc again and compare it with the previous commit. The MD5 has changed. If the client says “this version is identical to last week’s” but the hash is different, you have caught a silent change before it breaks the model.

Step 5: go back to the exact data behind an old model

Back to the question from Friday’s meeting. Find the production model’s commit (with git log) and check it out. In the example below, a1b2c3d is only an illustrative hash. Replace it with a real hash from your repo:

git checkout a1b2c3d
dvc checkout

One rule worth memorising: every time you run git checkout, also run dvc checkout. The reason goes back to the point made earlier. DVC does not keep history itself, Git does, and Git only knows about the .dvc file. So the first command only reverts the .dvc file. The second brings the data in your working directory back to the version that file points to.

How long you wait depends on how much data there is, so with large datasets, don’t promise the client a figure until you have measured it yourself.

Step 6: retrain only what needs retraining

For dvc repro to have anything to do, the repo needs a pipeline. Below is a minimal dvc.yaml with a single stage called filter, which reads the client data and writes out a filtered file. This is a simplified example: filter.py is a script you write yourself, and the full dvc.yaml syntax is in the DVC docs.

stages:
  filter:
    cmd: python filter.py
    deps:
      - data/data.xml
      - filter.py
    outs:
      - data/filtered.csv

DVC pipelines work like a build system. You run:

dvc repro

DVC reruns only the stages whose inputs have changed, and prints lines such as “Stage ‘filter’ didn’t change, skipping” for the rest. Try it now. Run dvc repro twice in a row and the second run will skip filter. Then dvc add a new data drop and run it again, and the stage will actually execute.

After each run, the hashes of any changed dependencies and outputs are written to dvc.lock. Commit this file to Git, because it pins down the pipeline state of a reproducible run. Once you are comfortable with one stage, add feature-building and training stages following the same pattern.

Common mistakes

The most common mistake is running only git checkout and then training, while the data/ folder still holds the latest version. The “old” model you think you have reproduced is really a new model trained with old code.

The second is committing the .dvc file but forgetting dvc push. Your colleagues or the client’s server get the commit but not the data that goes with it. Treat git commit and dvc push as an inseparable pair.

The third is not committing dvc.lock. You then have versioned data but cannot prove which training run used which inputs.

What does this skill look like at a client site?

In the field, DVC’s value lies not in the commands but in the answers they let you give. When a client asks why this week’s model differs from last week’s, you point to two commits, two MD5 hashes and the dvc.lock from each run. The argument moves from gut feeling to evidence.

The first thing to do on a new project: ask the client where data is allowed to be stored, and choose a remote accordingly before the first data drop arrives. Setting up the workflow after you already have three “final” versions is much harder.

For developers in Vietnam moving into FDE roles, look out for job descriptions that mention “reproducibility”, “data versioning” or “MLOps at client sites”. A CV line such as “versioned 12 client data drops with DVC on S3, able to reproduce any model from a commit” says more than a list of tools.

Clients will always send more data. The only question is whether, by the tenth drop, you still know which model learned from which version.

7 sources
Read next on the roadmap · Stage 5: DeploymentBuild CI/CD for ML models with CML: post metric comparisons on every pull requestWith one workflow file and three Python scripts, reviewers can see whether your change makes the model better or worse before they merge it.