# DVC in practice: make every new client data drop a traceable commit

> When a client asks "what data was last month's model trained on?", an FDE should answer with a commit hash, not from memory.

Bản gốc: https://fdetimes.net/en/guides/dvc-version-client-data-batches/

On Monday the client sends this week's data file. On Wednesday they send a "fixed" version. By Friday the new model is performing worse, and the first question in the meeting is: which version was the production model trained on?

If your answer is "probably Wednesday's", you have just lost some of the client's trust. FDEs run into this all the time when deploying ML in the field. Data keeps changing, and while the code is in Git, the data usually sits in a folder as something like `final_v3_moi.csv` ("final v3 new").

This guide walks through a workflow you can set up on a laptop in one evening using DVC. The DVC homepage sums up its philosophy as managing data the same way people manage code, and the Get Started docs call it "Git for data" outright.

## What will you build, and what do you need first?

Picture a retail client that sends a `data/data.xml` file every week. The exercise has three goals: each data drop maps to a Git commit, the actual data lives in storage the client has approved, and you can get back any drop whenever you need it.

You need Git, DVC installed, and a Git repo initialised for DVC as described on the Get Started page. One point to understand before typing any commands: technically, DVC itself is not a version control system. It leaves the history to Git, and the contents of the `.dvc` file determine which version of the data is in effect.

## Steps 1–2: turn the first data drop into a commit

Add the file to DVC:

```bash
dvc add data/data.xml
```

DVC creates a metadata file, `data/data.xml.dvc`. Open it and take a look. It records the size, the number of files and, most importantly, an MD5 hash. That hash identifies the data version.

Then commit the metadata file to Git, not the data itself:

```bash
git add data/data.xml.dvc
git commit -m "Dữ liệu khách: đợt tuần 1"
```

(The commit message reads "Client data: week 1 batch".) From this point on, the data's metadata sits in Git history right next to the source code. The example above is trimmed down. When you run `dvc add`, read DVC's output, because it may suggest other files to `git add`.

## Step 3: put the data where the client allows

If Git holds only a small file, where does the actual data go? DVC uses remotes: external storage for data and models, with support for S3, Azure, GCS, SSH and many others. This matters to FDEs because the client usually decides where its data is allowed to live.

```bash
dvc remote add myremote s3://mybucket
dvc push
```

`dvc push` uploads the data in the cache to the configured remote. This is a simplified example. Before you configure a real remote for a client, read the Remote Storage page in the DVC docs carefully.

To check that it worked, open the bucket. The data you just pushed should be there. If the client only allows an internal server over SSH, you change the remote URL, not the workflow.

## Step 4: the second data drop arrives

The client sends a new version. You overwrite the file and repeat the same routine:

```bash
dvc add data/data.xml
git add data/data.xml.dvc
git commit -m "Dữ liệu khách: đợt tuần 2"
dvc push
```

Open `data/data.xml.dvc` again and compare it with the previous commit. The MD5 has changed. If the client says "this version is identical to last week's" but the hash is different, you have caught a silent change before it breaks the model.

**Điểm mấu chốt:** The hash in the .dvc file is evidence. An engineer's memory is not.

## Step 5: go back to the exact data behind an old model

Back to the question from Friday's meeting. Find the production model's commit (with `git log`) and check it out. In the example below, `a1b2c3d` is only an illustrative hash. Replace it with a real hash from your repo:

```bash
git checkout a1b2c3d
dvc checkout
```

One rule worth memorising: every time you run `git checkout`, also run `dvc checkout`. The reason goes back to the point made earlier. DVC does not keep history itself, Git does, and Git only knows about the `.dvc` file. So the first command only reverts the `.dvc` file. The second brings the data in your working directory back to the version that file points to.

How long you wait depends on how much data there is, so with large datasets, don't promise the client a figure until you have measured it yourself.

## Step 6: retrain only what needs retraining

For `dvc repro` to have anything to do, the repo needs a pipeline. Below is a minimal `dvc.yaml` with a single stage called `filter`, which reads the client data and writes out a filtered file. This is a simplified example: `filter.py` is a script you write yourself, and the full `dvc.yaml` syntax is in the DVC docs.

```yaml
stages:
filter:
cmd: python filter.py
deps:
- data/data.xml
- filter.py
outs:
- data/filtered.csv
```

DVC pipelines work like a build system. You run:

```bash
dvc repro
```

DVC reruns only the stages whose inputs have changed, and prints lines such as "Stage 'filter' didn't change, skipping" for the rest. Try it now. Run `dvc repro` twice in a row and the second run will skip `filter`. Then `dvc add` a new data drop and run it again, and the stage will actually execute.

After each run, the hashes of any changed dependencies and outputs are written to `dvc.lock`. Commit this file to Git, because it pins down the pipeline state of a reproducible run. Once you are comfortable with one stage, add feature-building and training stages following the same pattern.

## Common mistakes

The most common mistake is running only `git checkout` and then training, while the `data/` folder still holds the latest version. The "old" model you think you have reproduced is really a new model trained with old code.

The second is committing the `.dvc` file but forgetting `dvc push`. Your colleagues or the client's server get the commit but not the data that goes with it. Treat `git commit` and `dvc push` as an inseparable pair.

The third is not committing `dvc.lock`. You then have versioned data but cannot prove which training run used which inputs.

## What does this skill look like at a client site?

In the field, DVC's value lies not in the commands but in the answers they let you give. When a client asks why this week's model differs from last week's, you point to two commits, two MD5 hashes and the `dvc.lock` from each run. The argument moves from gut feeling to evidence.

The first thing to do on a new project: ask the client where data is allowed to be stored, and choose a remote accordingly before the first data drop arrives. Setting up the workflow after you already have three "final" versions is much harder.

For developers in Vietnam moving into FDE roles, look out for job descriptions that mention "reproducibility", "data versioning" or "MLOps at client sites". A CV line such as "versioned 12 client data drops with DVC on S3, able to reproduce any model from a commit" says more than a list of tools.

Clients will always send more data. The only question is whether, by the tenth drop, you still know which model learned from which version.

**Thử ngay tuần này:**

- Take a public dataset, split it into 3 simulated 'deliveries' and commit each one with dvc add + git commit.
- Check out the first delivery, run dvc checkout and confirm that the MD5 in the .dvc file matches the commit.
- Add one specific line to your CV: how many data drops you versioned, which type of remote you used and how you reproduced an old training run.

## Nguồn

- [DVC](https://dvc.org/)

- [Get Started with DVC](https://doc.dvc.org/start)

- [Data Versioning (DVC Get Started)](https://doc.dvc.org/start/data-management/data-versioning)

- [Remote Storage (DVC User Guide)](https://doc.dvc.org/user-guide/data-management/remote-storage)

- [The Complete Guide to Data Version Control With DVC](https://www.datacamp.com/tutorial/data-version-control-dvc)

- [dvc repro (DVC Command Reference)](https://doc.dvc.org/command-reference/repro)

- [Git](https://git-scm.com/)
