# “How much data do we lose if it goes down?” Answer with RPO, RTO and a real restore

> “Don’t worry, we have daily backups” is not an answer. It is a number you haven’t worked out, and the client will work it out for you on the worst possible day.

Original: https://fdetimes.net/en/guides/rpo-rto-how-much-data-will-we-lose/

The demo has just ended, and the client’s head of operations asks what sounds like a simple question: “If the system goes down, how much data do we lose?” Many engineers reflexively reply, “Don’t worry, we have daily backups.” It sounds reassuring, but it dodges the question.

The client is not asking whether you have backups. They are asking for a number. Give them two well-grounded numbers, state their limits plainly, and you will earn more trust than any architecture slide. Dodge, and they will discover the real number themselves on the worst day.

For an FDE, this is a skill that follows you through every deployment. The systems you build on site touch the client’s real data, so sooner or later the question comes, usually from the person who signs the contract.

## One question that is really two

AWS Well-Architected defines RTO (Recovery Time Objective) as the maximum acceptable delay between the interruption of service and its restoration. RPO (Recovery Point Objective) is the maximum acceptable amount of time since the last data recovery point.

Google Cloud’s DR guide puts RPO more plainly: it is the maximum length of time during which data might be lost in a major incident.

An easy way to remember it: RTO counts forward from the outage to when the service is back. RPO counts backwards to the last good copy. “How much data do we lose?” is an RPO question, but the person asking almost always cares about RTO too, even if they have not said so.

AWS asks for both numbers to be set for each workload, not as one figure for the whole company. The rest of this piece shows how much that detail matters.

## If you back up at 2am, what is your RPO?

Imagine the client is a retail chain. Its orders database is snapshotted every day at 02:00. If the disk fails at 23:00, the best copy you have is the 02:00 one, so you lose 21 hours of orders.

Now take the worst case: the incident happens at 01:59, one minute before the next snapshot. You lose almost a full 24 hours. The real RPO of this system is 24 hours, and that is the number to give the client.

The AWS DR whitepaper says outright that backup frequency determines the recovery point you can achieve. Advisera, an ISO 27001 training provider, offers a handy rule: if your RPO is four hours, you must back up at least every four hours. If the client can tolerate losing at most one hour of orders, daily snapshots are not enough; you need far more frequent backups.

RTO has to be measured; it cannot be guessed. Restoring a large database, checking integrity, repointing the application, flushing caches: each step eats time in ways only a real run will reveal.

## Why you must never promise “no data loss”

At this point many people think straight away of continuous replication to bring RPO to zero. AWS acknowledges that continuous replication gives a near-zero backup lag. But it may not protect you against data corruption or deliberate deletion.

The reason is simple. A faulty migration script drops the `orders` table, and the replica follows suit within seconds.

AWS writes that this holds even for a multi-site active/active setup: when data is corrupted, deleted or obfuscated to the point of being unusable, recovery time is always greater than zero and the recovery point is always some point before the corruption was detected.

AWS goes further: even with every best practice in place, RTO and RPO remain above zero, meaning some loss of availability and data.

Well-Architected lists choosing unrealistic targets such as “zero data loss” as an anti-pattern. The same list includes setting targets so aggressive that costs rise while the business has no need for them.

## Ask first, propose later

A good FDE does not open with a solution. AWS suggests some discovery questions worth borrowing: what is the maximum amount of data that can be lost before the business suffers unacceptable impact? Can that data be recreated from another source?

The second question often changes the whole design. If online orders can be reconciled against the payment gateway, the RPO for the orders table may be relaxed. By contrast, consultation notes typed by hand by staff are gone for good once lost.

Then tier the workloads. AWS recommends grouping workloads by business impact into critical, high, medium and low, each with its own RTO/RPO. The table below is a hypothetical example for the retailer above, purely to illustrate the thinking:

| Workload (hypothetical) | Tier | Proposed RPO | How to achieve it |
|---|---|---|---|
| Payments, orders | Critical | A few minutes | Continuous log backup + periodic snapshots |
| Inventory | High | 1 hour | Back up at least hourly |
| BI reports | Medium | 24 hours | Daily snapshots, rebuild from source |
| AI pipeline logs | Low | Can be lost | Re-run from source data |

RTO determines the DR model. AWS draws a clear distinction: pilot light cannot serve requests without additional action, while warm standby takes traffic immediately, albeit at reduced capacity. For the medium tier, pilot light or restoring from backup is enough.

Once you have the numbers, your answer in the meeting room might sound like this:

> “For orders, we lose at most a few minutes of data if the infrastructure fails, and the system is back within the time we measured during our drill. If data is deleted by mistake, we restore to the most recent copy before the moment it was detected. That is a limit of every system, and we have alerting to shorten the detection window.”

## Backups on paper and GitLab’s 18 hours

GitLab’s 2017 postmortem is a lesson many people who run systems still cite. GitLab.com was down for about 18 hours, and the GitLab team itself wrote that they had no idea their backups were failing until it was too late. The backups existed, but only on paper.

**Key point:** A backup that has never been test-restored is not a recovery plan.

Google Cloud recommends exactly what GitLab skipped: once a DR plan is written, test it regularly, record every issue that comes up and adjust the plan accordingly. For an FDE, a restore drill is the only way to make the RTO you quote to a client a measured number.

On your next deployment, you can work in this order. First, list every data store in the deployment, including the vector stores and object storage that people often forget.

Next, sit down with the client to tier the workloads and agree RTO/RPO in writing. Then align the backup schedule with the RPO, run a real restore, time it, and set up alerts for when a backup job fails.

Four mistakes come up again and again, and each one creates a promise you cannot keep:

- Talking about “having backups” instead of talking about RPO.
- Treating a replica as a backup.
- Setting one number for every system.
- Never test-restoring.

If you are looking to move into an FDE role, you can put this skill on your CV with a concrete line: “Designed backups for a 1-hour RPO, ran quarterly restore drills, measured RTO of X minutes.” When reading job descriptions, look out for phrases such as “disaster recovery”, “on-call” or “customer-facing incidents”.

When you see them, prepare an answer to the very question that head of operations asked, in case the interview turns to the topic.

The next time a client asks “how much do we lose if it goes down?”, you will already have the answer from your latest drill, with the run time and the measured number.

**Try this week:**

- Pick a database you operate, note when its backups run, and calculate the worst-case RPO: an incident one minute before the next backup.
- Restore the latest backup into an isolated environment, time it from start until the application can read the data, and record that number in the runbook.
- Draft a three-sentence answer to “how much data do we lose if it goes down?” covering RPO, RTO and the limit for accidental deletion.

## Sources

- [REL13-BP01 Define recovery objectives for downtime and data loss](https://docs.aws.amazon.com/wellarchitected/latest/reliability-pillar/rel_planning_for_recovery_objective_defined_recovery.html)

- [Disaster recovery planning guide](https://docs.cloud.google.com/architecture/dr-scenarios-planning-guide)

- [Disaster recovery options in the cloud](https://docs.aws.amazon.com/whitepapers/latest/disaster-recovery-workloads-on-aws/disaster-recovery-options-in-the-cloud.html)

- [RTO and RPO: What is the difference between Recovery Time Objective and Recovery Point Objective?](https://advisera.com/27001academy/knowledgebase/what-is-the-difference-between-recovery-time-objective-rto-and-recovery-point-objective-rpo/)

- [Postmortem of database outage of January 31](https://about.gitlab.com/blog/postmortem-of-database-outage-of-january-31/)
