FDE PulseFDE jobs open 441New in 7 days 29Companies hiring 47Remote-friendly 24%Median US pay $216kTop hirer Databricks 125
VI

The newspaper of the Forward Deployed Engineer

Guides

Hands-on: separating dev, staging and prod for a pipeline when the client only has real data

When a client has a single database full of personal information, your job is to build staging without letting any sensitive data leave production.

In brief

  • Do not copy production data straight into staging: personally identifiable information (PII) goes with it.
  • Mask before data leaves production, then build staging and dev from the masked copy.
  • Describe all three environments in code and promote them through Git, so they stay identical apart from parameters.
ShareLinkedInFacebookX
GraphicFrom one real database to three safe environments
  1. 1Classify columnsMark PII columns with the client; do not rely on automatic PII detection
  2. 2Define environments as codeOne shared module; each of dev, staging and prod has only its own parameter file
  3. 3Mask inside productionCreate a masked view with keyed hashing so PII never leaves production infrastructure
  4. 4Create staging and devBranch or snapshot from the masked copy; refresh by recreating
  5. 5Deploy through GitEach environment points to a folder; changes to prod must go through a pull request

PII stays only in production; every other environment is built from the masked copy and deployed through Git.

Graphic: FDE Times

Picture your first week on a client site. You ask where staging is and are told there is only one database. It is also production, and it holds the name, phone number and address of every customer. The pipeline you are about to write has to run on that data, but you cannot experiment on it.

The usual first instinct is to “copy production to another database to be safe”. That instinct is wrong. Neon, the serverless Postgres provider, warns plainly that production data cannot be safely copied to staging or test environments, because personal information travels with the copy.

This guide walks through five steps to build three environments from scratch. All three are described in code, data is masked before it leaves production, and deployment goes through Git. The code here is simplified illustration to show the structure. It is not configuration you can run as is.

What will you build, and what do you need?

The running example is a hypothetical delivery company with an orders table. Its columns are id, customer_name, phone, address, order_value and created_at. The end result is three environments, dev, staging and prod, sharing one definition, with dev and staging seeing only masked data.

You need read access to the production schema, an infrastructure-as-code tool and a Git repository. On tooling: OpenTofu works for both cloud and on-premises. AWS CDK lets you write infrastructure in a programming language and deploy it through CloudFormation. On GCP, Infrastructure Manager runs Terraform configuration as a managed service. If the client uses Kubernetes, add Argo CD.

Step 1: Mark every table that holds PII, now

Before writing a line of infrastructure, sit down with someone who understands the client’s data and classify every column. In the example, customer_name, phone and address are PII. order_value and created_at must stay intact, because your pipeline computes on exactly these columns.

Do not skip this step. Neon states clearly that it does not detect PII automatically and masks nothing unless the user declares rules. So rather than relying on an automatic detection tool, treat your classification file as the primary input for every masking rule.

Check: you have a classification file in which every column belongs to one of two groups: keep as is, or mask. The client has reviewed and agreed to it.

Step 2: One definition, three sets of parameters

Environment as Code means describing an entire environment in code, including services, configuration and dependencies. According to Bunnyshell, this makes each environment a replica of production and so limits the “works on my machine” problem. For that to hold, the definition must be written only once.

# Illustrative repo layout, not tied to any specific tool
pipeline-infra/
  modules/pipeline/     # single definition: DB, job, scheduler
  envs/
    dev.yaml            # parameters only
    staging.yaml
    prod.yaml

With CDK, you get the parameters, conditionals and loops of a real programming language. The snippet below is a sketch in plain Python, not the CDK API:

# Sketch: same function, different parameters
ENVS = {
    "dev":     {"data_source": "masked", "size": "small"},
    "staging": {"data_source": "masked", "size": "medium"},
    "prod":    {"data_source": "live",   "size": "large"},
}
for name, params in ENVS.items():
    build_pipeline(name, **params)   # a function you write yourself

Check: comparing the parameter files tells you nothing, because by design they contain only parameters. Compare what gets generated instead: the resource list after rendering the full configuration for staging and prod, or the list of resources actually deployed. If they differ in anything other than size and data source, such as a service that exists only in prod, staging is no longer a true replica.

Step 3: Where you mask is the most important question

There are two places to put the masking step, and their risks are very different. The Database Lab Engine (Postgres.ai) documentation describes both.

Mask before leaving production

  • PII stays only in production
  • Non-prod environments receive only the masked copy
  • Masking rules run on the client’s infrastructure

Mask in the non-prod environment

  • PII is physically copied out of production
  • Real data sits on non-prod infrastructure
  • One wrong access permission is enough to leak data

On a client site, choose the left column. A simple approach is to create a masked view inside production and grant the export process read access to that view only:

-- Illustrative, Postgres with the pgcrypto extension;
-- the secret key lives only in production and never travels with the export
CREATE VIEW orders_masked AS
SELECT id,
       'customer_' || id AS customer_name,
       encode(hmac(phone, current_setting('app.mask_key'), 'sha256'), 'hex') AS phone,
       'redacted'        AS address,
       order_value,
       created_at
FROM orders;

Pay attention to the choice of masking function. Phone numbers have a small value space, so if you hash with plain md5(phone) and no key, anyone holding the export can hash every possible phone number and match them back.

Keyed hashing (HMAC) preserves uniqueness, so the pipeline can still join or count returning customers, but without the key it cannot be reversed. order_value stays intact so that results computed in staging match reality.

Check: run a query over the export that searches for strings shaped like phone numbers. The result must be zero. Keep this query and run it on every data refresh.

Step 4: Staging and dev both come from the masked copy

Once the masked copy exists, every non-prod environment draws its data from it. Neon describes a branching approach: child branches inherit the masked data, and refreshing staging is simply a matter of recreating the branch. If the client does not use Neon, the same principle applies with snapshots of the masked copy.

The practical gain is that each engineer’s dev environment can be deleted and recreated at any time without anyone touching production. A bug that only appears with real data can now be reproduced on data of the same shape, just without real names.

Check: delete your dev branch or database and recreate it. If it takes more than a few minutes, or needs someone to grant you access to production, the process still has a gap.

Step 5: Promote environments through Git, not by hand

The final step keeps the three environments from drifting apart over time. Argo CD is a GitOps continuous delivery tool for Kubernetes. Its principle is that application definitions, configuration and environments should all be declarative and under version control. Argo CD can also manage multiple clusters.

# Illustrative: each environment points to a folder in the repo
deploy/
  dev/       -> dev cluster
  staging/   -> staging cluster
  prod/      -> prod cluster

To promote a change from staging to prod, open a pull request that edits the prod/ folder rather than running commands by hand on your own machine. For non-Kubernetes infrastructure, OpenTofu aims at the same thing: a consistent workflow for managing infrastructure throughout its lifecycle.

Check: pick any change running in prod and trace it back to its commit. If you cannot find one, someone is making manual edits.

Three common mistakes on client sites

The first mistake is masking with a script on an engineer’s laptop. The real data has already reached the laptop, which puts you squarely in the right-hand column of the comparison, only on a machine that is even harder to control than a non-prod server.

The second is forgetting new columns. Imagine the client adds an email column next month and the masking rules are not updated. With a tool that masks only by manually declared rules, such as Neon, that column goes straight through to staging.

The phone-number scan from step 3 will not catch this, because it does not look at the schema.

The fix is a separate check, call it a schema check: read the current column list of the table in production, compare it with the classification file from step 1, and fail when it finds an unclassified column.

Run it alongside the scan on every data refresh.

The third mistake is letting staging drift from prod because of “just a quick fix”. After a few weeks, staging no longer proves anything about prod.

How does this skill come up in FDE interviews?

When reading FDE job descriptions, watch for phrases such as “deploy in customer environments”, “regulated data” or “on-prem”. When you see them, prepare an answer on how you handle real data, because the role will probably put you in exactly the situation this guide describes.

On your CV, do not just write “used Terraform, Argo CD”. Write in terms of outcomes, for example: built three environments from a single definition, masked PII inside production, with automated checks running on every data refresh. Those three points show you understand the risk a client holding real data carries.

At your first session with a client, do not open by asking for production admin access. Ask to sit down and classify columns with the person who owns their data. It is the cheapest step, and a good way to start building trust.

Was this article useful?

Use with your AI assistantAsk Claude ↗Ask ChatGPT ↗
7 sources
Read next on the roadmap · Stage 5: DeploymentHands-on: build a lightweight SNS–SQS–Lambda fanout pipeline in a client's AWS accountOne topic, three queues, one Lambda. This exercise fits within the Free Tier and teaches the duplicate and lost-message failures you will meet in production on a client site.