# AWS Glue, Azure Data Factory and Google Dataflow: how to read the ETL pipeline a client already runs

> On a client site, an FDE rarely gets to choose the ETL tool. Usually the job is to take over a pipeline that is already running, so the first task is to understand it before changing anything.

Bản gốc: https://fdetimes.net/en/tools/aws-glue-azure-data-factory-google-dataflow/

"Are you on Azure Data Factory or Data Factory in Microsoft Fabric?" At a client running Azure, this should be one of the first questions an FDE asks.

The reason is on Microsoft's own overview page. Microsoft still describes ADF as a managed cloud service for complex ETL, ELT and hybrid data integration projects. On the same page, however, it calls Data Factory in Fabric the "next generation" of ADF and points new users to Fabric.

The question reflects a wider reality: when FDEs work with data, they rarely start from scratch. By the time you arrive, the client is usually already running AWS Glue, Azure Data Factory or Google Dataflow. The data your agent or model needs to read flows through those pipelines.

You do not need to be an expert in all three. You do need to be able to read the client's pipeline within your first week.

## Three services, three vocabularies

AWS describes Glue as a serverless data integration service for discovering, preparing, moving and integrating data from multiple sources. Crawlers sit at the centre of Glue: a crawler infers the schema and writes it to the Glue Data Catalog. Glue Studio lets users build transformation workflows by drag and drop and run them on a serverless ETL engine based on Apache Spark.

Those who prefer to write code can use Spark, Python or Scala.

ADF has its own set of concepts: pipelines, activities, datasets, linked services, data flows and integration runtimes. Mapping data flows run on a Spark cluster that starts when needed and shuts down afterwards, so users do not manage the cluster. The integration runtime is the bridge between activities and linked services. In other words, it decides where the pipeline actually runs.

Google defines Dataflow as a unified batch and stream processing service that runs at scale. Dataflow was announced in June 2014, entered public beta in April 2015 and is now built on the open-source Apache Beam project. As load changes, the service adds or shuts down worker VMs automatically.

## One problem, three places to look first

Picture a retail chain that wants an agent to answer questions about inventory. Every night, order files land in a data lake, while the inventory table sits in a database in the office. Your job is to bring both sources into one clean place the agent can query. The table below suggests where to look on each platform.

| Step | AWS Glue | Azure Data Factory | Google Dataflow |
|---|---|---|---|
| Know what the data contains | Check the schema the crawler wrote to the Data Catalog | Read the datasets and linked services | Read the source-reading step in the Beam code |
| Reach the on-premises database | Open the job, note the connection to the office database, then ask the network team about its route and access permissions | Check which integration runtime acts as the bridge | Find the database-reading step in the Beam code, then ask the network team whether the worker VMs can reach the office database |
| Transform | Glue Studio or PySpark jobs | Mapping data flows running on Spark | Beam transforms |
| When it runs | On a schedule, on demand or on an event | Open the pipeline, look at the order of activities and ask the client what triggers it | Ask the client how the job is launched, and whether it is a batch or streaming job |

Read across the rows and the hardest question in this example sits in the second one. The integration runtime is the bridge between activities and linked services, so on Azure it is the first place to check whether the pipeline can connect to the office database.

Before writing any transform, sit down with the client's network and security teams to pin down that bridge.

**Điểm mấu chốt:** Do not change a client's pipeline until you know where it gets its schema and where it runs.

## Where are the limits?

All three services hide much of the infrastructure, and that is exactly what makes FDEs complacent. Spark starting and stopping on its own, or workers scaling automatically, does not mean the transformation logic is correct. A schema inferred by a crawler from the data is still a guess, so check it with whoever owns that data source.

Azure adds a strategic question. Microsoft says existing ADF workloads can be upgraded to Fabric. Any proposal to build something new on ADF should therefore come with a direct question to the client: do they plan to move to Fabric? Skip it, and you may find yourself planning an upgrade just after handover.

With Dataflow, the difficulty lies in the programming model rather than the infrastructure. Beam uses one model for both batch and streaming. The model is powerful, but people used to writing sequential scripts will need time to adjust.

## What to learn first

If you have only a month, learn Spark first. On Glue, you write ETL jobs directly in Spark. In ADF, Spark sits underneath mapping data flows and is managed by the service, so Spark knowledge helps you understand how data flows run rather than being something you operate yourself.

Then learn enough Beam to read a Dataflow pipeline. Finally, practise bringing pipelines into a release process: ADF fully supports CI/CD through Azure DevOps and GitHub, and enterprise clients will ask about it.

When reading FDE job descriptions, note which ETL services are mentioned to decide which platform to brush up on first. On your CV, do not write "knows AWS Glue". Write that you read an existing pipeline, found where the schema had drifted and shipped the fix through CI/CD.

That sentence shows a hiring manager exactly the work an FDE does on a client site.

The client's ETL tool was almost certainly chosen before you arrived. What you control is how quickly you understand it.

**Thử ngay tuần này:**

- Create an AWS free tier account, point a crawler at a folder of CSV files on S3, then look at the schema it inferred in the Data Catalog
- Write an Apache Beam pipeline that reads a file and counts lines by key, and run it locally before thinking about running it on Dataflow
- Take an FDE job description, underline the ETL services it names, then write one CV line about how you read, fixed or brought such a pipeline into CI/CD

## Nguồn

- [Introduction to Azure Data Factory - Azure Data Factory | Microsoft Learn](https://learn.microsoft.com/en-us/azure/data-factory/introduction)

- [What is AWS Glue? - AWS Glue Developer Guide](https://docs.aws.amazon.com/glue/latest/dg/what-is-glue.html)

- [AWS Glue - Serverless Data Integration Service](https://aws.amazon.com/glue/)

- [Dataflow overview | Google Cloud Documentation](https://docs.cloud.google.com/dataflow/docs/overview)

- [Google Cloud Dataflow - Wikipedia](https://en.wikipedia.org/wiki/Google_Cloud_Dataflow)
