ISO 20022 for FDEs: check the payment data before bringing AI into a bank
A bank that has moved to the new message standard does not necessarily have clean data. Measure how clean it is before a model learns the wrong things.
In brief
- "We're on ISO 20022" tells you only the message format. Legacy address data entered as free text is still often incomplete, abbreviated and inconsistent.
- Swift and the EPC have both pushed back their deadlines for banning fully unstructured addresses. For some time yet, data will mix all three forms: structured, hybrid and unstructured.
- An FDE's first job is to profile the data source by source and record its lineage. Only after that should anyone talk about models.
You have just started work at a bank. The brief is to build an agent that helps the compliance team read international payment messages, group counterparties by city and country, and flag anomalies. At the first meeting, the client’s technology lead says with confidence: “We’re on ISO 20022 now. All our data is structured.”
That is only half true. The other half tends to surface about three weeks after the model goes live, usually as a list of counterparties that makes no sense and that nobody can explain.
Your job in the first week is to find that other half before it finds you: what ISO 20022 actually delivers, where the data is still dirty, and how dirty it is.
What does ISO 20022 promise?
Legacy MT messages pack a lot of information into lines of free text. ISO 20022 replaces that with clearly defined fields for names, postal addresses and remittance information. A machine can read each part separately instead of guessing where to split a string.
For cross-border payments on Swift CBPR+, the period in which MT and ISO 20022 coexisted officially ended on 22 November 2025.
In Vietnam, the ACH run by NAPAS has implemented ISO 20022, BIDV has used the standard since 2023, and in November 2025 the State Bank of Vietnam held a two-day conference in Hanoi to push its adoption in interbank payments.
So at a Vietnamese bank, “have you moved to the standard?” is no longer the question worth asking. The right question is: which internal systems are still feeding old free-text data into the new messages?
FTI Consulting argues that because ISO 20022 is a structured format, both classical ML and generative AI become much more effective. That is the view of a consultancy, not the result of independent measurement. The argument is reasonable, but it holds only when the data actually sits in the structured fields.
Why doesn’t “we’re on ISO 20022” mean the data is clean?
A new message format does not fix existing data. Addresses typed as free text over decades are often inconsistent, abbreviated or incomplete. One risk to check as soon as you start profiling is whether that legacy data has simply been carried over into the free-text address lines of the new messages.
That is why a single table can contain three forms. The structured form breaks out each component: street name, building number, town, country. The hybrid form has some components in fields while the rest remain in free-text lines.
The fully unstructured form consists only of text lines. EPC rules allow hybrid as a transitional step, while fully unstructured addresses are banned. The problem is that the date the ban takes effect keeps moving.
Under the original timetable, Swift CBPR+ would stop accepting fully unstructured addresses from 14 November 2026. Swift has postponed that date and expects to publish a new timeline before December 2026. In the SEPA area, the EPC has also postponed its 15 November 2026 deadline and will set a new one later.
In practice, the client’s data will mix all three forms for some time. Before putting any date into a plan, check Swift’s SR2026 page.
A first-day profiling exercise
Picture the same counterparty appearing in two records. The XML below is for illustration only:
<!-- Hoàn toàn phi cấu trúc -->
<PstlAdr>
<AdrLine>12 NGUYEN HUE Q1 HCMC VN</AdrLine>
</PstlAdr>
<!-- Có cấu trúc -->
<PstlAdr>
<StrtNm>Nguyen Hue</StrtNm>
<BldgNb>12</BldgNb>
<TwnNm>Ho Chi Minh City</TwnNm>
<Ctry>VN</Ctry>
</PstlAdr>
The first is fully unstructured; the second is structured. A person sees at a glance that both are the same address. A pipeline that groups by the town field sees only the second record, because the first has no town. So the first thing to write is a small classification script:
from collections import Counter
def classify(pstl_adr):
# Không có thẻ, hoặc có thẻ <PstlAdr/> rỗng: đều là thiếu địa chỉ
if pstl_adr is None or len(pstl_adr) == 0:
return "missing"
tags = [c.tag.split("}")[-1] for c in pstl_adr]
has_lines = "AdrLine" in tags
has_fields = any(t != "AdrLine" for t in tags)
if has_fields and not has_lines:
return "structured"
if has_fields and has_lines:
return "hybrid"
return "unstructured"
def profile(records):
# records: list of (source_system, pstl_adr_element)
return Counter((src, classify(adr)) for src, adr in records)
The detail to watch is the empty tag. The comment in classify makes the point: a missing tag and an empty <PstlAdr/> both count as a missing address. Without that check, an empty <PstlAdr/> with no child elements would be counted as “unstructured”, and you would wrongly tell the client that a source holds free-text addresses when in fact it holds no address at all.
Suppose you sample 1,000 records and find 380 structured, 270 hybrid and 350 unstructured. If the agent reads only the town field, and only some of the hybrid records have that field, then at best it sees 650 records.
At least 350 counterparties, more than a third of this hypothetical sample, drop out of the analysis without a single error message.
The total matters less than the figures by source. If nearly all 350 unstructured records come from one legacy core system, you have a concrete question for the client: will data from this source be standardised, or excluded from scope?
Five steps before touching the model
Step one is to sample by source and by time period rather than drawing randomly from the whole table, because old and new data often look very different. Step two is to run the profiling above and present the result as a table of source × address form. That table lets business owners understand the problem without reading XML.
Step three is to ask about data governance. FTI stresses that a metadata catalogue, field lineage and retention policies are preconditions for using ISO 20022 data in a way that is reliable and auditable. At a bank, the question “where did this field come from, and who transformed it?” will come up in an audit sooner or later.
Step four is to identify which data has real value for the problem. According to FTI, richer data on the ultimate debtor and ultimate creditor helps banks understand their customer base better. But you can use it only if those fields are actually populated, so measure their fill rate the same way you measured addresses.
Only at step five do you reach the model. For hybrid and unstructured data, decide explicitly whether to parse it, flag it or exclude it, and record that decision so the compliance team can review it.
Three mistakes new FDEs often make
The most common mistake is taking “we’re on the standard” at face value without opening the data. The second is using an LLM to parse free-text addresses into fields and then overwriting the original data.
That destroys lineage. When an auditor asks, nobody can show which values came from the customer and which the machine inferred.
The third is building a plan around a fixed deadline. As noted above, these dates are still moving, so a roadmap stating “no unstructured addresses after November 2026” may be wrong on the day it is presented.
What does this skill look like on a CV?
When reading job descriptions for FDE or solutions engineer roles in banking, look for the keywords ISO 20022, pacs.008, CBPR+, NAPAS or data lineage. For these roles, be ready to show that you can both write code and talk to payment operations teams.
On your CV, do not write “knowledge of ISO 20022”. Write a line with numbers: how many records you profiled, what share of unstructured addresses you found in which source, and how you decided to handle them.
If you have no real project yet, a small repository using synthetic data with a README explaining the three address forms is enough to open a good interview.
On a banking project, an FDE’s value is not in getting a model into production. It is in pointing out, like the 350 records in that hypothetical sample, the data that is quietly disappearing, before the model learns from whatever is left.
5 sources
- CBPR+ is live: what ISO 20022 means in practice (Mambu) · 2026-07-27
- The November 2026 Structured Address Deadline: What Every PSP Needs to Do Now (Clearing Post) · 2026-03-11
- From Compliance to Competition: Unlocking the Benefits of ISO 20022 (FTI Consulting) · 2026-02-18
- Conference promotes ISO 20022 in interbank payment system (Viet Nam News) · 2025-11-17
- BIDV adopts ISO20022 for cross-border payments in its payment hub system · 2023-10-10