Kubeflow or SageMaker, Vertex AI, Azure ML: ask who will run the pipeline before comparing features
Microsoft lists Kubeflow as an equivalent of its own product, and Google runs KFP code directly. A feature comparison no longer helps an FDE pick a platform.
In brief
- Pipeline code is now largely shared: Vertex Pipelines runs KFP code, and Microsoft itself lists Kubeflow Pipelines as the open-source equivalent of Azure ML pipelines.
- The real difference is who operates it: KFP needs a customer team that owns Kubernetes, while Vertex and SageMaker take on the orchestration infrastructure.
- Ask first who will be on call when the pipeline breaks, then choose the tool and interface that suit that person.
In its official documentation on machine learning pipelines, Microsoft draws up its own table and places Kubeflow Pipelines in the “open-source equivalent” cell for Azure Machine Learning pipelines. Both serve the same persona, the data scientist, and cover the same stretch from data to model.
A cloud provider is calling an open-source project the equivalent of its own product. Google, for its part, lets Vertex Pipelines run code written with the KFP SDK directly. When the feature boundaries are this blurred, comparing features is close to useless, and the deciding question moves elsewhere: once you leave the customer’s site, who will operate the pipeline?
For an FDE this is not a theoretical question. You are often the one who builds the first pipeline, but not the one who lives with it. Pick a platform that does not suit the people taking over, and however elegant the pipeline is, it will slowly rot.
Pipeline code now speaks almost one language
Start with Kubeflow. According to its official repository, Kubeflow Pipelines are reusable end-to-end ML workflows, written with the KFP SDK and run on Kubernetes. You declare each step, wire the output of one step into the input of the next, then compile the whole thing into a pipeline definition.
The crucial point is that this definition is not locked to your cluster. When Google announced general availability of Vertex Pipelines in November 2021, it said plainly that the service supports two open-source libraries, Kubeflow Pipelines and TensorFlow Extended. A pipeline written in KFP therefore has a path to Google’s managed service without its logic being rewritten.
Azure has not stayed outside this trend. Microsoft lets you build pipelines through the CLI, the Python SDK or the Designer UI, and in the same comparison table, Azure ML pipelines and Kubeflow Pipelines share a row, both emphasising distribution, caching, code-first development and reuse. The vendors are using a common vocabulary.
Even the most common distinction has lost its edge. A blog post from JFrog, a vendor with its own perspective, sums it up this way: Kubeflow focuses on orchestration and pipelines, while SageMaker leans towards data science. That is a difference in product emphasis, not in whether the customer’s pipeline will run.
The real difference: who is on call when the pipeline breaks?
Reread the documentation, this time paying attention only to what it says about operators. The Kubeflow repository states that KFP can be installed as part of Kubeflow Platform or deployed as a standalone service. Either way the consequence is the same: someone on the customer’s side has to own the Kubernetes cluster, upgrade it and receive the alerts when it misbehaves.
The cloud providers say the opposite, and say it bluntly. Google writes that Vertex AI handles provisioning and scaling the infrastructure that runs pipelines, that customers pay only for the resources used while the pipeline runs, and that data scientists get to focus on ML.
AWS describes SageMaker Pipelines as a serverless orchestration service built specifically for MLOps and LLMOps, meaning the customer does not have to run the orchestration layer themselves.
The second promise lies in the interface. SageMaker Pipelines has a drag-and-drop interface in SageMaker Studio, and Azure ML has the Designer UI alongside its SDK and CLI. The vendors are designing for people who do not want to read YAML, not just for platform engineers.
The Azure ML documentation also contains a line worth pinning to the meeting-room wall: data engineers, data scientists and ML engineers each own their own steps. A pipeline has no single owner. So the question of “who operates it” has to be split into three layers: who writes the steps, who fixes the steps, and who keeps the machinery underneath alive.
Picture two customers. Customer A is an e-commerce company with three data scientists, running on AWS, with nobody who knows Kubernetes well. Customer B is a bank whose platform team has run Kubernetes for years and wants everything on its own infrastructure.
On features, either could use any of the options. On operators, A should almost certainly go with SageMaker Pipelines, because nobody on the team can carry a cluster. B has a legitimate reason to deploy KFP standalone, because its platform team is already used to being on call for exactly that kind of system.
AWS advertises that SageMaker Pipelines can run tens of thousands of concurrent ML workflows in production. An impressive number, but it settles nothing for customer A, which probably runs only a few pipelines a day. The more useful number to ask about is how many people on the customer’s side can read a pod’s logs at midnight.
Four options, ranked by who operates them
Put the documentation together and the picture becomes clearer when the “who operates it” column comes first instead of the feature column:
| Option | Who operates it, per the documentation | Ways of building pipelines mentioned | Best suited to |
|---|---|---|---|
| Kubeflow Pipelines (standalone or within Kubeflow Platform) | The customer’s team, on their own Kubernetes | KFP SDK | Teams that already have a platform group comfortable with Kubernetes |
| Vertex Pipelines | Google provisions and scales the infrastructure | KFP SDK or TFX | GCP users who want to keep KFP code without the burden of a cluster |
| SageMaker Pipelines | AWS, as serverless orchestration | Drag-and-drop interface in SageMaker Studio | AWS users with small teams less familiar with infrastructure code |
| Azure ML pipelines | Data engineers, data scientists and ML engineers each own their own steps | CLI, Python SDK or Designer UI | Azure users where several roles share ownership of the steps |
Look at the last column and you will see that the decision is really made before any feature page is opened. The cloud the customer already uses and the team the customer already has rule out almost every option.
What does a pipeline that is “ready to move” look like?
There is one move that reduces risk when you are not yet sure who will operate the pipeline: write it with the KFP SDK and keep the logic free of any single cloud’s proprietary APIs. The illustrative code below has just two steps (the identifiers are Vietnamese: lam_sach means “clean”, huan_luyen means “train”, du_lieu_vao means “input data”):
from kfp import dsl, compiler
@dsl.component(base_image="python:3.11")
def lam_sach(du_lieu_vao: str) -> str:
# chỉ xử lý dữ liệu, không gọi API riêng của cloud nào
return du_lieu_vao
@dsl.component(base_image="python:3.11")
def huan_luyen(du_lieu: str) -> str:
return "model-uri"
@dsl.pipeline(name="churn-pipeline")
def churn(du_lieu_vao: str):
sach = lam_sach(du_lieu_vao=du_lieu_vao)
huan_luyen(du_lieu=sach.output)
compiler.Compiler().compile(churn, "churn.yaml")
The comment in the first component reads: “only processes data, does not call any cloud-specific API”.
The generated churn.yaml file can run on a KFP deployment that the customer’s platform team runs themselves, and because Vertex Pipelines supports KFP, it also has a path to Google’s managed service. If, six months later, the customer decides it no longer wants to be on call for a cluster, you change where it runs, not the logic.
To keep that advantage, separate everything tied to the environment, such as storage paths or project names, into parameters passed into the pipeline rather than hard-coding them in components. In your first week on a customer site, the first thing to do is write on the whiteboard the name of the person who will receive the alert when the pipeline fails, and only then open the IDE.
Reading job descriptions to see who carries the pager
The same skill helps you choose jobs. A job description that mentions Kubernetes, Helm and Kubeflow usually signals that you or the customer will be the operator, so be ready to talk about cluster upgrades and incident handling.
One that stresses SageMaker Studio, Designer or working with stakeholders suggests you will hand over to data scientists, and the key skill is designing so that others can fix things.
For developers in Vietnam looking to move into FDE roles, the cheapest way to practise is to run the same KFP file in two places: a local cluster and Vertex Pipelines. On your CV, do not just list “Kubeflow, SageMaker”. Spell out that you chose a platform based on which team would operate it, and whom you handed it over to.
In interviews, when asked which tool to choose, try answering with a question of your own: after the FDE leaves, who gets woken up when the pipeline breaks? Practise that reflex in advance, because the tool can change after a single meeting, while the operating team is what you have to design around.
5 sources
- GitHub - kubeflow/pipelines: Machine Learning Pipelines for Kubeflow
- Announcing Vertex Pipelines general availability · 2021-11-11
- Amazon SageMaker Pipelines
- What are machine learning pipelines? - Azure Machine Learning | Microsoft Learn · 2026-09-09
- A Brief Comparison of Kubeflow vs. SageMaker · 2022-11-10