FDE PulseFDE jobs open 434New in the last 7 days 27
VI

The newspaper of the Forward Deployed Engineer

Guides

Encryption, tokenization or masking: choosing the right tool for each customer data field

Security reviews rarely turn on algorithms. The client's security team will ask three things: who holds the keys, who can see what, and whether one person's data can be deleted when they ask.

Một thanh dữ liệu màu kem đi vào nút chia màu cam rồi tách thành ba nhánh, kết thúc lần lượt ở một tấm vé màu xanh ngọc, một thanh bị che nửa bằng các chấm tròn và một ổ khoá.

In brief

  • Each data field needs its own technique. Encryption answers who holds the key, tokenization answers who holds the mapping table, and masking answers who can see what.
  • Envelope encryption with KMS separates keys from data and logs every use of a key. Have both ready to show in the review.
  • Avoid shuffling and blanket nulling. The pipeline must also be able to delete one person's data, even after it has been masked or tokenized.
ShareLinkedInFacebookX
GraphicField by field: how to protect it, how to delete it
Protection techniqueOn a deletion request
Customer IDTokenization, mapping table kept in a separate vaultFind tokens by owner, delete downstream rows, then delete from the vault
Full nameDynamic masking based on the viewer's permissionsOnly one copy of the original exists, so deleting it in the main DB is enough
Phone numberEnvelope encryption with KMSDelete the ciphertext; earlier decryptions remain traceable in CloudTrail
Card number in ticket textDetected and replaced before sending to the LLMThe card number never reached the LLM, so nothing needs deleting on the model side
Incident description (dev copy)Static masking, identifier column kept as a tokenFind rows via the customer ID token and delete them, with no need to recover data

A protection technique is only right if every field still has a path to deleting one person's data.

Graphic: FDE Times

Picture your second week at a client. The pipeline that feeds customer support tickets into an LLM for classification is working, and the demo went smoothly. Then the client’s security team sends over a questionnaire. The first question reads: “Which systems does a customer’s phone number pass through, who can read it, and where are the keys?”

Many engineers answer by reflex: “We encrypt everything.” That answers almost nothing. The security team needs to know how each field is handled, who holds the keys and who can see what.

For an FDE, answering these questions clearly should be treated as part of the job, not something left to the security team. You understand the pipeline better than anyone, so you are also the best person to explain it.

This article covers three tools: encryption, tokenization and data masking. The running example is the support ticket pipeline above, taken one field at a time. It ends with a list of the mistakes that tend to add weeks to a security review.

Each tool protects data in a different way

Encryption keeps the data intact, readable only by whoever holds the key. So the most important question about encryption is where the key is kept. OWASP recommends storing encryption keys separately from the encrypted data where possible. The reason is simple: if the key sits next to the data, whoever steals the disk gets both.

Masking goes the other way. Wikipedia defines it as modifying sensitive data so that it has little or no value to anyone without authorised access. AWS puts it in terms closer to engineering work: masking creates a fake version of the data with the same structure as the real thing, so developers and analysts can still work with data that looks real.

There are two kinds of masking. Static masking is applied to a stored copy of the database, usually one used for dev or test environments. Dynamic masking is applied at query time.

AWS describes it as a filter: the original data stays unchanged, and what is displayed changes according to the permissions of the person viewing it. This fits well with role-based access control in the client’s systems.

Security tokenization sits between the two. In the design below, the real value is replaced with a random token, and the table mapping tokens back to real values is kept somewhere separate and tightly protected. Downstream systems only see the token, but they can still join and count on it.

The two meanings of “tokenization”

This is where AI people often get crossed wires in meetings with security teams. NVIDIA explains that tokens in AI are small units of data produced by breaking a larger body of information into pieces. Converting input data into tokens before a model processes it is also called tokenization.

Tokenization in this sense protects nothing. A phone number that passes through an LLM tokenizer is still a phone number, just cut into several pieces. When a security team asks “do you tokenize PII?”, they mean the security sense.

If your pipeline involves both meanings, use two different names in the documentation, such as “LLM tokenizer” and “PII token vault”, so nobody confuses them.

Field by field through the support ticket pipeline

Suppose each ticket has five fields: customer ID, full name, phone number, a bank card number the customer accidentally pasted into the text, and the incident description. Some of the company’s customers live in California.

The California Attorney General’s office lists several rights under the CCPA, including the right to request deletion of personal information and the right to limit the use and disclosure of sensitive personal information. IBM notes that the CCPA requires reasonable safeguards, and that the fine for an intentional violation is $7,500.

These two rights shape the design directly. The right to delete means the pipeline must be able to find and delete one person’s data. The right to limit means you have to decide which fields may pass on to downstream systems or to the LLM. Applying both requirements to each field gives the following table:

Field Technique Reason
Customer ID Tokenization Downstream systems still need to join; to delete, remove the mapping row
Full name Dynamic masking Support staff see the real name, analysts only see a masked one
Phone number Envelope encryption Staff sometimes need to call the customer back, so it must be reversible
Card number in ticket text Detect and replace before sending to the LLM The model does not need the card number to classify the incident
Incident description Static masking in the dev copy Engineers need realistic-looking data to debug

For the phone number, the standard implementation is envelope encryption. According to the AWS KMS documentation, the data key is generated on an HSM and protected by a KMS key. The data key encrypts the data, and the KMS key (acting as the KEK) wraps the data key. This matches OWASP’s recommendation to keep keys separate from data.

import os, boto3
from cryptography.hazmat.primitives.ciphers.aead import AESGCM

kms = boto3.client("kms")

def encrypt_field(plaintext: bytes, key_id: str) -> dict:
    dk = kms.generate_data_key(KeyId=key_id, KeySpec="AES_256")
    nonce = os.urandom(12)
    ct = AESGCM(dk["Plaintext"]).encrypt(nonce, plaintext, None)
    # Store only the KMS-wrapped data key, never the plaintext key
    return {"ct": ct, "nonce": nonce, "edk": dk["CiphertextBlob"]}

This approach also gives you something to show in the review. According to AWS, every request to a customer managed key is recorded as a CloudTrail event. When the security team asks “who decrypted phone numbers last week?”, you answer with logs, not promises.

For the customer ID, the key point in designing the token vault is to tie each token to an owner. That way, when a deletion request arrives, you can find all of that person’s tokens before deleting them.

import secrets

def tokenize(owner_id: str, value: str, vault) -> str:
    token = "tok_" + secrets.token_hex(16)
    vault.put(token, owner=owner_id, blob=encrypt_field(value.encode(), KEY_ID))
    return token

def forget(owner_id: str, vault, stores) -> None:
    tokens = vault.tokens_of(owner_id)
    for store in stores:  # main DB, masked dev copy, cache
        store.delete_where_token_in(tokens)
    vault.delete_by_owner(owner_id)  # any leftover tokens now point nowhere

The order inside forget is deliberate. The statically masked dev copy holds no real names or phone numbers, but its customer ID column is kept as tokens rather than masked.

So you can still find the rows to delete in the dev copy, while the masked columns remain unrecoverable: getting from a token back to a real person means going through the vault, which is tightly protected.

Five steps before the security review

The first step is to build a classification table for every field, as above, including the ones you think are harmless. Free text such as the incident description often contains things nobody expected. The second step is to map the data flow: where each field appears in its original form and where it appears in processed form.

The third step is to choose where the keys and the token vault live, separate from the main data store. Ideally use a KMS managed by the client, so that the power to revoke keys stays in their hands.

The fourth step is to be able to prove it with logs: pull a few sample CloudTrail events ready to present. The last step is to rehearse a deletion request end to end, covering the main database, the masked dev copy (finding rows via the customer ID token), caches and every log containing prompts sent to the LLM.

Mistakes that cost you the security team’s trust

The first mistake is masking by shuffling, that is, swapping values between rows. Wikipedia notes that this can be reversed if someone works out the shuffling algorithm. Masked data must not be recoverable.

The second mistake is nulling out an entire column to save time. It is simple, but it damages data integrity and also reveals that the column has been masked. A dev copy full of nulls also leaves engineers unable to debug anything, and they end up asking for access to the real data.

The third mistake is keeping keys next to data, for example a plaintext data key in a config file in the same repo. The fourth is masking the identifier column in the dev copy as well, so that when a deletion request arrives nobody knows which rows belong to whom.

The fix is to keep that column as tokens, as described above: sensitive columns are still masked irreversibly, and only the link to the owner goes through the vault.

If you are preparing to apply for FDE roles, read job descriptions closely for keywords such as “PII”, “KMS”, “data residency” or “security review”. On your CV, instead of writing “experienced in data security”, write a specific line such as: “Designed envelope encryption and a token vault for a ticket pipeline, passed the client’s security review, supported per-customer data deletion.”

Plenty of engineers know encryption algorithms. What security teams need is someone who can answer three questions for every field: who holds the key, who sees what, and whether deleting one person really deletes everything.

7 sources
Read next on the roadmap · Stage 5: DeploymentL4 and L7 load balancers and reverse proxies: running your service behind a customer's networkWhen your service has to run behind a customer's load balancer, three faults tend to surface together: users' real IPs disappear, sessions jump between servers and health checks report the wrong thing. All three can be fixed once you know what each network layer can see.