Customer data onboarding is the FDE's core job, not a chore before the "real work"
No model gets anywhere if the equipment codes in two systems don't match. The person who fixes that is usually the FDE.
In brief
- Tagup lists ETL and integration as hard skills for FDEs. Palantir asks for experience with, or interest in, large-scale data.
- Enterprise AI pilots often fail because data sits in many separate systems that are hard to integrate. FDEs have to close that gap.
- The process has five steps: inventory sources, profile, map keys, reconcile, hand over.
- 1Inventory sourcesList systems, owners, export methods and how often the data is updated
- 2ProfileCount nulls, check distributions, LEFT JOIN to find unmatched records
- 3Map keysNormalize codes and time zones, export an exceptions file for the customer to confirm
- 4ReconcileEach run prints match rate against real sensors and checks joins don't duplicate rows
- 5Hand overDocument every rule and its reason so whoever takes over doesn't have to guess
Profile and reconcile first, write the pipeline later. Every mismatched row must be confirmed by the customer.
Graphic: FDE Times
In your first week on site, you open an export from the customer’s maintenance system. The equipment code column says “PUMP-07”. The sensor database says “P7”. The demo is booked for Friday morning, and no code will run until something recognises those two names as the same pump.
Many engineers treat this stage as “cleanup” to get through before the interesting part. Job descriptions disagree. Tagup’s FDE posting spells out one of the role’s tasks: bringing in the imperfect data that the customer’s real operations still depend on. The same posting lists Python or TypeScript, SQL, ETL and integration as hard skills.
Paraform, citing MIT NANDA research via The New Stack, says enterprise AI pilots fail largely because company data is scattered across many separate systems that are hard to integrate.
The Pulse newsletter from The Pragmatic Engineer describes today’s FDE role as heavy on integration and close customer support, and light on building new systems from scratch. If that description holds, data onboarding is a skill to practise before you apply.
Why does data expose the real problem?
Palantir created the role in the early 2010s and calls these engineers Delta internally. According to the company’s job posting, FDSEs work directly with customers to quickly understand their biggest problems. They work in small teams with little oversight and own critical projects end to end.
If you own a project from start to finish, the first thing you touch is almost always the customer’s data. In a small team with little oversight, there probably won’t be a separate data engineering team to do that part for you.
The same posting also asks for experience with, or interest in, using large-scale data to solve valuable business problems.
There is a deeper reason too. Mismatched equipment codes often mean that two departments have never had to share data. Working out why a join fails therefore tells you something about how the customer’s organisation works.
Paraform describes the FDE as a hybrid of software engineering and customer deployment. Data onboarding is where those two halves meet.
An example: two systems, one pump
Imagine the customer is a water utility that wants to predict when its pumps will fail. There are two data sources. The file work_orders.csv from the maintenance system has asset_code, opened_at in local time, and a failure_type column typed in by hand by technicians. The sensor_devices table in Postgres has device_id, with vibration readings recorded in UTC.
Profiling comes first. The pipeline can wait. A simple SQL query shows how far apart the two sources are (the alias order_count means “number of work orders”):
-- How many maintenance asset codes have no match on the sensor side?
SELECT w.asset_code, COUNT(*) AS order_count
FROM work_orders w
LEFT JOIN sensor_devices s ON s.device_id = w.asset_code
WHERE s.device_id IS NULL
GROUP BY w.asset_code
ORDER BY order_count DESC;
The comment asks: how many maintenance codes have no match on the sensor side? Suppose almost none of them match. That doesn’t mean the data is dirty. The two systems just name things differently. The next step is to normalise both sides to a shared key, convert the timestamps to one time zone, and join again:
import re
import pandas as pd
def normalize(code) -> str | None:
m = re.search(r"(\d+)$", str(code).strip().upper())
return f"PUMP-{int(m.group(1)):02d}" if m else None
wo = pd.read_csv("work_orders.csv")
wo["asset_key"] = wo["asset_code"].map(normalize)
wo["opened_at_utc"] = (
pd.to_datetime(wo["opened_at"])
.dt.tz_localize("Asia/Ho_Chi_Minh")
.dt.tz_convert("UTC")
)
# conn: connection to the customer's Postgres
sensors = pd.read_sql("SELECT device_id FROM sensor_devices", conn)
sensors["asset_key"] = sensors["device_id"].map(normalize)
merged = wo.merge(sensors[["asset_key", "device_id"]],
on="asset_key", how="left", indicator=True)
unmatched = merged[merged["_merge"] == "left_only"]
unmatched.to_csv("needs_customer_confirmation.csv", index=False)
The important part is the last two lines. The rule “take the number at the end of the code” is only your hypothesis. Work orders that can’t be linked to a sensor must go to the customer’s owner of that data for confirmation, not be quietly dropped. Here that file is needs_customer_confirmation.csv, meaning “needs customer confirmation”. (The comment above conn notes that it is the connection to the customer’s Postgres.) The exceptions file often starts the most useful conversation of the week.
The last step is reconciliation, which runs every time the pipeline runs. Here total is the total and matched the number matched. The assertion messages read “two sensors merged into one key” and “the join duplicated work orders”, and the printed line reports how many work orders were linked to a real sensor:
total = len(wo)
assert sensors["asset_key"].dropna().is_unique, "Two sensors collapsed into the same key"
assert len(merged) == total, "The join duplicated work orders"
matched = int((merged["_merge"] == "both").sum())
print(f"{matched}/{total} work orders linked to a real sensor ({matched/total:.1%})")
The two assertions catch two ways the normalisation rule can go wrong without anyone seeing it: two pumps get merged into one, or a work order is duplicated by the join. The match rate is measured against the real list of device_id values from the sensor side, not against your own normalised output.
You take that printed line into the customer meeting. It turns “the data is a bit messy” into progress you can measure.
A process you can run yourself
Start with an inventory: which sources exist, who owns each one, how the data is exported and how often it is updated. Then profile each source: count nulls, look at value distributions, try joining the keys. Only once you know where the mismatches are should you define a canonical schema and a shared key.
Next, write the mapping rules as explicit code, with an exceptions file and someone on the customer side who signs off on it. Every pipeline run must print its reconciliation results, so the match rate is always measured on the latest data.
At handover, the documentation should spell out each rule and why it exists, so whoever takes over doesn’t have to guess.
Mistakes that quietly kill projects
| Common mistake | Consequence | What to do instead |
|---|---|---|
| Writing the pipeline before profiling | Key mismatches surface just before the demo | Run a LEFT JOIN to count orphaned records from day one |
| Quietly dropping unmatched rows | The model trains on a skewed dataset and nobody knows | Export an exceptions file and send it to the customer to confirm |
| Guessing what hand-typed columns mean | Failure types are misclassified | Ask the technicians who entered the data |
| Mixing local time with UTC | Failure events drift away from the sensor signal | Normalise time zones as soon as the data is read |
| No reconciliation numbers | No answer to “is the data good enough yet?” | Print the match rate against the real sensor list and check the row count after the join on every run |
How do you show this skill when applying?
When you read a job description, look for phrases such as ETL, integration, “imperfect data” or “large scale data”. They tell you where your time will go. If a JD talks only about models and says nothing about customer data, ask about it in the interview.
On a CV, a line such as “unified data from three systems, raising the key match rate to a level the customer accepted within two weeks” carries far more weight than “proficient in pandas”. You can also add a small repo to your portfolio: two mismatched public datasets, an exceptions file and a reconciliation step.
A repo like that is concrete evidence that you have done exactly the kind of work these JDs describe.
On demo day, the customer won’t remember which algorithm your model used. They will remember the first time they saw the maintenance tickets and sensor signals for the same pump on one screen.
Was this article useful?
Thanks for the feedback!