FDE PulseFDE jobs open 441New in 7 days 29Companies hiring 47Remote-friendly 24%Median US pay $216kTop hirer Databricks 125
VI

The newspaper of the Forward Deployed Engineer

Guides

Go-live cutover: moving the data, running in parallel and planning the way back

A go-live night at a client can fall apart even when the code is right, because nobody knows who has the authority to say "roll back", or which data to roll back to.

Go-live cutover: moving the data, running in parallel and planning the way back
Photo: Paul Rivenberg and Mary Pat McNally, MIT Plasma Science and Fusion Center / CC BY 3.0

In brief

  • A cutover is a fixed sequence, not a button, and the final backup is also your rollback source.
  • Rollback is simple only while the new system has not taken real data. Once new transactions exist, you need a way to get that data back into the old system.
  • Every plan needs one named person who decides on rollback, a bounded time window and at least one rehearsal.
ShareLinkedInFacebookX
Timeline of the cutover night from 22:00 to 00:45. At 22:00 writes to the old system are frozen; at 22:15 the final backup, already restore-tested; at 22:45 the final sync; at 23:30 routing is switched to the new system (the highlighted milestone); at 23:45 testing and reconciliation within a 60-minute limit; at 00:45 the client's project lead decides whether to fix forward or roll back. Before 23:30 is the easy rollback zone using the 22:15 backup; after 23:30 is the hard rollback zone, where fail-forward DB, dual write or backup/restore must be chosen in advance.
The 23:30 routing switch is the line: from then on the new system takes live orders, and the 22:15 backup is no longer enough to go back.

Picture this: it is 10pm on a Friday. The client’s old order system is still running, and the pipeline you have spent three months building is ready. You have four hours to make the switch. The most important question of the night has nothing to do with code. If something goes wrong at 1am, who decides to roll back, and what data do you roll back to?

For an FDE, go-live night is when the client really judges you. A demo shows them what the system can do. A cutover tells them whether they can trust you with their operations. That makes it a skill worth practising before you are the one holding the runbook on a real night.

A cutover is a sequence, not a button

AWS Prescriptive Guidance describes cutover as a fixed sequence: freeze writes to the old system, take a final backup, run a final data sync, switch routing, then test.

The order matters. You freeze first so the old system takes no new transactions while you are copying. Otherwise the final sync will be missing some orders and nobody will notice.

The final backup is not just for the archive. AWS states plainly that this same backup can be used for an emergency rollback. So record exactly when the backup was taken and do a test restore before cutover night, to be sure it actually works when you need it.

Next you choose between moving everything at once (big bang) and moving in phases. According to AWS, a phased approach is more complex and takes longer, but usually means less downtime and faster rollback.

That is why it is the usual choice for business-critical production systems. When a client has thousands of users, you might move one branch or 5% of traffic first, then expand.

Parallel runs: let the machine compare before you trust it

One way to test a new system before trusting it is to run it alongside the old one on real traffic. GitHub’s Scientist library does this at the code level: on every call, both branches run in a randomised order, their results are compared, mismatches are recorded, and the user still gets the result from the old branch.

For data, Stripe describes four steps: dual-write to the old and new tables, move reads to the new table, move writes, and only then delete the old data. Existing data is backfilled into the new table.

If you are replacing a rule-based order classifier with a new model, this pattern applies directly: run the model in shadow mode and log every result that differs from the old rules so you can review them before switch-over day.

Parallel running has its limits, though. DualEntry’s runbook documentation, written for accounting systems, warns that double data entry doubles the workload, and that the second copy is the one most likely to drift.

They recommend a clean break at an accounting period boundary, with the old system set to read-only. The two views point to one principle: letting machines run in parallel and compare results is fine; making the client’s staff enter data twice should be avoided.

Rollback is easy only before new data arrives

This is where many plans fall short. AWS notes that if you roll back after the new system has taken real transactions, you may also have to restore that data to the old system.

There are three options: a fail-forward database, dual writes, or backup and restore. The longer the new system runs on real data, the harder the way back becomes.

According to AWS, a rollback plan needs three things: trigger checkpoints, a data strategy, and a named person who decides whether to fix forward or roll back.

A post on the AWS blog adds a time-boxing rule. If, within the agreed window, the team cannot clearly state what the problem is and how to fix it, roll back.

AWS calls for the rollback procedure to be written into the cutover plan itself and rehearsed. The AWS blog post goes further and compares it to a fire drill: it has to be done regularly so the procedure is not forgotten.

A sample runbook for go-live night

AWS recommends a runbook that sets out the start time, end time, order and owner of each task, together with a RACI matrix. Imagine a distributor moving its order intake to a new system you are deploying. A condensed runbook might look like this:

Time Task Owner Decision point
22:00 Freeze writes to the old system Client DBA Confirm no new orders are coming in
22:15 Final backup, record the checkpoint Client DBA Backup restores successfully in a test
22:45 Final data sync FDE Record counts match on both sides
23:30 Switch routing to the new system DevOps From here the new system takes real orders
23:45 Testing and reconciliation FDE + business team 60-minute limit; roll back when time is up
00:45 Go-live decision Client project lead Fix forward or roll back

The last column is the most important. Suppose the old system shows 12,480 orders for the day and the new one counts 12,477. Those three missing orders are a trigger checkpoint: either you find the cause within the agreed window, or you roll back.

But roll back with what? From 23:30, once routing has switched, the new system may already have taken real orders, and the 22:15 backup does not contain them.

Restoring the backup alone would lose every order that came in after 23:30, so the plan must choose one of the three data strategies above in advance, such as dual-writing back to the old system throughout the testing window.

DualEntry offers a rule worth copying into every runbook: do not declare go-live until every line on the reconciliation checklist passes. In the example above, the checklist might include: order counts match, total order value matches, and 50 randomly sampled records are identical in both systems.

Five mistakes that drag go-live night into the morning

The first is having no named person to decide on rollback. Five engineers stare at the same mismatched number, each wants another ten minutes to try a fix, and the time limit stops meaning anything.

The second is having a written rollback plan that nobody has run, so you only discover when you need it that a restore takes three hours, not thirty minutes. The third is forgetting that once routing switches, the new system starts taking real data, while the rollback plan still assumes nothing has happened.

The fourth is making client staff enter data in parallel “just to be safe”. The fifth is declaring go-live before reconciliation is finished, simply because everyone is tired.

If you are preparing to move into an FDE role, look for phrases such as “migration”, “go-live support” or “deployment ownership” in job descriptions. On a CV, a line like “wrote and coordinated the cutover runbook, rehearsed rollback on staging, went live with no data loss” tells a recruiter more than any list of frameworks.

Clients may forget which model you used. They will remember whether, on go-live night, someone could answer the question “what if it breaks?”

Was this article useful?

Use with your AI assistantAsk Claude ↗Ask ChatGPT ↗
6 sources
Read next on the roadmap · Stage 5: DeploymentA client sends a security questionnaire: how FDEs use SOC 2 and ISO 27001 to get through reviewSometimes the thing blocking a project is not the hardest code but a spreadsheet with hundreds of rows. One compliance consultancy estimates that sending a SOC 2 report first can cut the 260-question CAIQ to roughly 78 to 104 questions you have to answer yourself.