FDE PulseFDE jobs open 434New in the last 7 days 27
VI

The newspaper of the Forward Deployed Engineer

Tools

MLflow: answering “which model is running?” with runs, versions and aliases

At a customer site, a good model is not enough. You also have to show which version is running in production, where it came from and how to rebuild it.

Bốn chiếc hộp giống nhau xếp trên một kệ, chiếc vương miện treo trên chiếc hộp màu cam mới nhất, còn đường nét đứt phía trên cho thấy nó vừa được dời sang từ chiếc hộp thứ hai.

In brief

  • Each time you register a model under an existing name, MLflow increments the version number automatically.
  • An alias is a name that points to a specific version and can be moved to another, which makes it well suited to marking the version in production.
  • Every version links back to the run or notebook that produced it, so you can trace a production model back to its original experiment.
ShareLinkedInFacebookX
GraphicFrom experiment to running model in MLflow
  1. 1Log runsEach training run records parameters, metrics and artifacts to the dashboard
  2. 2Compare runsUse the dashboard to pick the run with the best results
  3. 3Register the modelRegister under the same name and the version number increments automatically
  4. 4Assign the champion aliasThe alias points to the running version and can be moved to a new one
  5. 5Trace backGo from the version to the original run or notebook to explain and reproduce the result

The alias tells you which version is running; the link back to the run tells you how it was made.

Graphic: FDE Times

In the MLflow Model Registry, each time you register a model under a name that already exists, the version number goes up automatically. It sounds like a small detail, but it settles the most awkward question in any ML project at a customer site: “Which model is actually running?”

For an FDE, that question tends to arrive at the worst possible moment. This week’s predictions are off, and the customer wants to know whether someone has just swapped the model. If the answer lives in one engineer’s memory, or in a filename like model_final_v2_fix.pkl, you are losing credibility.

MLflow is not the most glamorous tool you will learn. But it is the cheapest way to turn “I think it’s this one” into “it’s this one, produced by this run, with these parameters”.

Two things MLflow does: record and number

MLflow describes itself as the largest open-source AI engineering platform for agents, LLMs and ML models, released under the Apache 2.0 licence. For the purposes of this article, only two parts matter: Tracking and the Model Registry.

Tracking manages every experiment and its components. Each training session becomes a run, and you record its parameters and results so you can later compare runs against each other instead of digging back through notebooks.

The MLflow dashboard lists experiments, the runs inside them, and, for each run, its metrics, parameters and artifacts.

The Model Registry is the next step: a central store where several people can manage the full lifecycle of a model together. Three concepts matter here: auto-incrementing versions, aliases that point to a specific version, and links from each version back to the run, logged model or notebook that produced it.

An example: three versions, one “champion”

Imagine you are building a churn prediction model for a telecoms company. In the first week you run three experiments with three different parameter sets. Each run is logged to the dashboard with its evaluation metrics and model file.

You register all three under the same name, churn-model. The registry numbers them version 1, 2 and 3; you do not have to name them yourself. Version 2 performs best, so you give it the alias champion, and the customer’s serving system always loads the model through that alias rather than by version number.

The whole workflow fits into a few dozen lines of Python. The sketch below covers a single run, assuming you already have training and test data:

import mlflow
from mlflow import MlflowClient
from sklearn.linear_model import LogisticRegression
from sklearn.metrics import roc_auc_score

# 1. Log một run: tham số, metric và chính model
with mlflow.start_run():
    model = LogisticRegression(C=0.5).fit(X_train, y_train)
    auc = roc_auc_score(y_test, model.predict_proba(X_test)[:, 1])
    mlflow.log_param("C", 0.5)
    mlflow.log_metric("auc", auc)
    info = mlflow.sklearn.log_model(model, "model")

# 2. Đăng ký vào registry: cùng tên thì version tự tăng
mv = mlflow.register_model(info.model_uri, "churn-model")

# 3. Gán alias cho version được chọn
MlflowClient().set_registered_model_alias("churn-model", "champion", mv.version)

# 4. Phía serving chỉ biết alias, không biết số version
live_model = mlflow.pyfunc.load_model("models:/churn-model@champion")

(The comments read: 1. log a run with its parameters, metrics and the model itself; 2. register it, and the same name means the version increments; 3. assign the alias to the chosen version; 4. the serving side knows only the alias, not the version number.)

Two weeks later you retrain on fresh data and get version 4. After evaluating it, you simply repeat step 3 with version 4. The MLflow documentation calls aliases “mutable” references, meaning they can be pointed at a different version, so promotion is just moving a pointer, and the line of code in step 4 stays the same.

Now the customer asks: “Which model is running?” You open the registry: champion points to version 4, and version 4 links back to the run that produced it, which holds the full set of parameters and metrics. If you need to roll back, version 2 is still there, and the saved experiments let you reproduce the old results without guesswork.

When should an FDE bring MLflow to a customer?

Two properties make MLflow a good fit for deployment work. The first is neutrality: MLflow says it works with whatever cloud, framework or tools you already use. If the customer has already chosen its infrastructure, you do not have to persuade them to switch.

The second is that it can be self-hosted. You can run MLflow on your own servers and database, so there are no third-party limits or storage costs. For customers sensitive about their data, the ability to keep everything inside their own infrastructure may well be the deciding factor, so it is worth asking about explicitly during customer discovery.

The sensible moment to introduce MLflow is as soon as more than one person is training models on the project, or just before the first deployment. Waiting until something breaks to set up a registry means the old versions have already left no trace.

What MLflow will not do for you

MLflow only records what you log. If the pipeline does not record the data version or preprocessing steps in the run, the link from version to run leads to an incomplete picture, and you still cannot reproduce the result.

An alias is also just a pointer. Who is allowed to move champion, and what checks must pass before they do, is a process you have to agree with the customer. MLflow’s GitHub repository mentions evaluation features that help catch regressions before they reach production, but running them before every promotion is a matter of team discipline; the tool does not enforce it.

Self-hosting has its own cost, too. The MLflow server and database become one more system that someone has to operate, back up and manage permissions for.

What to learn first, and how to put it on your CV

The sensible order is tracking first, registry second. Get used to every training session becoming a run with parameters, metrics and artifacts, then learn to register versions and assign aliases. The LLM and agent features can wait until you are comfortable with the lifecycle of a classical model.

When reading job descriptions for FDE or ML engineer roles, look out for phrases such as “experiment tracking”, “model registry” and “reproducibility”. On your CV, do not just add “MLflow” to a list of skills.

Describe how you used aliases to mark the production version and traced an incident back to its original run. A sentence like that shows you understand the problem the tool solves, not just its name.

A good model wins you the demo. A clear registry is what keeps the customer’s trust the second time they ask, “What just changed?”

5 sources
Read next on the roadmap · Stage 5: DeploymentMichael Nygard's Release It!: the bedside book for FDEs who live in a client's productionA client API that suddenly hangs for 30 seconds per call can drain your service's threads in just 5 seconds. This 376-page book shows how to break that chain of failures at design time.