Why a backup restore test is the real control
Most teams can show that backups run. None of that tells you whether you can get your data back. Backups fail quietly: a job that skips a new database, an encryption key nobody can find, a restore that works but takes two days.
A backup restore test closes that gap. You take a real backup, restore it somewhere safe, and check that the data is complete and usable within the time your business needs. A disaster recovery (DR) test goes one level up: it checks whether you can bring a whole service back, with the right people, in the right order, after something big goes wrong.
Auditors know the difference. A screenshot of a successful backup job shows a backup exists. A dated restore test record shows the backup works.
Where backup and DR testing fits in SOC 2
Backup and recovery testing belongs mainly to the Availability category of the SOC 2 Trust Services Criteria. In general terms:
- A1.2 covers the protections and recovery infrastructure behind your availability commitments, including data backup and recovery processes.
- A1.3 covers testing your recovery plan procedures to show they support recovery.
- CC9.1, in the Common Criteria, covers identifying and mitigating the risk of business disruption. A business continuity and disaster recovery (BCDR) plan is the usual way to address it.
Availability is one of the four optional categories. You choose it if your customers rely on your service being up and your commitments to them say so. If you leave it out, A1.2 and A1.3 are not in your report, but CC9.1 still applies, so your auditor will usually still want a BCDR plan. Your auditor decides what evidence your scope needs. For how the categories fit together, see the SOC 2 readiness checklist.
RTO and RPO in plain terms
Every restore test and DR test is measured against two targets. Set them in your BCDR plan before you test, otherwise you have nothing to pass or fail.
- Recovery time objective (RTO): how long a system can be down before the damage is unacceptable. "We can restore the customer database within 4 hours."
- Recovery point objective (RPO): how much data you can afford to lose, measured in time. "We can lose at most 24 hours of data." Your backup frequency has to be at least as frequent as your RPO.
The targets differ by system. A small company might set targets like these (an example, not a recommendation):
| System | RTO | RPO | Why |
|---|---|---|---|
| Production database | 4 hours | 24 hours | Customers can't use the product without it |
| File storage for customer uploads | 8 hours | 24 hours | Needed, but the product works in a degraded mode |
| Internal wiki | 3 days | 7 days | Inconvenient to lose, not customer facing |
Pick numbers you can defend and actually meet.
How to run a quarterly backup restore test
Quarterly is a common cadence for small companies. It catches drift (new databases, changed jobs, rotated keys) without becoming a burden.
- Pick a system and a dataset. Rotate through your critical systems over the year, starting with the production database. Name the exact backup you will use: which system, which backup set, and its timestamp.
- Restore to an isolated environment. Use a separate account, network or project that can't reach production and that production can't reach. Lock down who can see it, because restored data is still customer data.
- Start the clock. Record when you began the restore. The time from "we decide to restore" to "the data is usable" is what you compare with the RTO.
- Verify integrity and completeness. Don't stop at "the restore command succeeded". Compare row counts or object counts with production at the backup's timestamp, check that recent records exist, run a checksum where your tools support it, and start the application against the restored data.
- Measure against RPO and RTO. The age of the newest data in the backup, compared with when the "failure" happened, is your achieved recovery point. The elapsed restore time is your achieved recovery time. Write down both, with the targets next to them.
- Record issues and follow-ups. Anything that went wrong or took too long gets a ticket with an owner: a missing table, a manual step nobody had documented, a credential that had expired.
- Clean up. Delete the restored copy and its environment, and note that you did. Leftover copies of production data are a confidentiality problem of their own.
- Get it reviewed. Someone other than the person who ran the test reads the record and signs off. That review is part of the control.
Restore test record template
Copy this table into your evidence page for each test. The example entries are illustrative.
| Field | Example entry |
|---|---|
| Period and test date | 2026 Q4, tested 2026-11-12 |
| Performed by | Platform engineer on rotation |
| Reviewed by | Engineering lead |
| System and dataset | Production database, all schemas |
| Backup used | Nightly backup, taken 2026-11-12 02:00 UTC |
| Restore target | Isolated test account, no network path to production |
| RPO target and achieved | Target 24 hours; newest record 8 hours before the test start |
| RTO target and achieved | Target 4 hours; usable after 1 hour 40 minutes |
| Integrity checks | Row counts within expected range on 12 key tables; application started and a test login worked |
| Result | Pass, Partial or Fail |
| Issues and follow-up tickets | Restore runbook missing one step; ticket raised and assigned |
| Cleanup confirmed | Restored copy and test environment deleted 2026-11-12 |
Attach the supporting material: restore logs or command output, the query results you compared, and screenshots with the date visible.
How to run an annual disaster recovery test
A DR test checks whether you can recover a whole service: infrastructure, data, configuration, secrets, the people and the decisions. Once a year is a common cadence. There are three broad ways to run one.
| Type | What happens | Effort and risk |
|---|---|---|
| Tabletop | The team walks through a scenario using the BCDR plan, step by step, without touching systems | Low. Finds gaps in the plan, contacts and decisions, but proves nothing about timing |
| Partial failover | You rebuild one critical service, or the whole stack in a separate environment, from backups and code | Medium. Gives real recovery times without affecting customers |
| Full failover | You move production to your recovery environment and run from it | High. The strongest evidence, and a real outage risk if it goes wrong |
For a first test, a tabletop followed by a partial failover of your most important service is a sensible mix. The tabletop format is the same one described in how to run an incident response tabletop exercise.
Plan the test
- Scenario. Make it specific: "Our primary cloud region is unavailable for 24 hours" or "A bad migration corrupted the production database and the corruption reached last night's backup."
- Roles. Name a test lead, the people who perform the recovery, a scribe who timestamps every step, and a decision maker who would declare a disaster in real life.
- Success criteria. Write them before you start: service restored within the RTO, data loss within the RPO, the plan followed without undocumented steps, and customers notified on the path the plan describes.
Run it and write it up
Follow the BCDR plan as written, not as people remember it. Note where the plan is wrong and carry on. Afterwards, compare the achieved times with the targets, list every issue with an owner, and update the plan.
Evidence to keep for each test
Auditors sample the periods in your observation window, so each test needs its own dated record. For more on organizing it, see SOC 2 evidence collection.
- Each quarterly restore test: the completed record above, the restore logs or output, the integrity check results, the issues raised, and the reviewer's sign-off with a date.
- Each annual DR test: the test plan (scenario, scope, roles, success criteria), the timestamped log of what happened, the RTO and RPO achieved against the targets, the issues and follow-up tickets, and the plan changes that came out of it.
- The documents behind them: an approved Backup Policy (what is backed up, how often, how long it is kept, how restores are tested) and an approved BCDR plan with the RTO and RPO targets.
Common mistakes
- Only checking that the backup job succeeded. A green job is not a restore. Restore the data and look at it.
- Restoring to production. Restoring over live data is how a test becomes an incident.
- Never testing the database. Teams restore a file or two because it is easy and skip the database because it is hard. The database is usually the system that matters most.
- No targets. Without an RTO and RPO, "it took six hours" is neither a pass nor a fail.
- Losing the follow-ups. Issues found in a test and never fixed show up again next quarter, and auditors notice.
Scheduling restore and DR tests in Confluence
Compliance in a Box, an app for Confluence Cloud, includes both tests as SOC 2 recurring activities when you select the Availability category: Backup Restore Test (quarterly, mapped to A1.2) and Business Continuity / DR Test (annual, mapped to A1.3). Availability also adds the Backup Policy. The Business Continuity & Disaster Recovery Plan is part of Security as well, so it is always provisioned and is mapped to A1.2, A1.3 and CC9.1. If you start with Security only, you can add Availability later and generate again; the app creates only the pages that are missing.
- An evidence page for every period. Each activity has its own page with the procedure and the evidence to keep. Before each period is due, the app creates an evidence page under it with a checklist: restore performed and integrity verified for the restore test, and test plan, RTO/RPO achieved and issues for the DR test. Paste in the record template and attach your logs.
- A schedule that matches your year. Periods follow your fiscal year and are due on their last day. Evidence pages open 14 days ahead by default; a Compliance Admin can change that with Edit schedule, for example to give a DR test more lead time.
- An owner and approvers. Compliance Admins assign them under Settings > Owners & approvers. Only the owner and Compliance Admins can edit the evidence pages. The owner is reminded with a Confluence task when the evidence page opens, 3 days before the due date and weekly once it is overdue.
- Version-bound approval. The owner selects Submit for approval from the Compliance byline item on the evidence page, and the approval is tied to that exact page version. Self-approval is allowed but flagged, so choose a reviewer other than the person who ran the test.
The Backup Policy and the BCDR plan are policy pages with their own approvals and a 12-month review. When a DR test changes your recovery targets, edit the plan and submit the new version for approval. The details are in the Recurring activities chapter and the SOC 2 content reference, and the SOC 2 compliance calendar shows where both tests sit among your other recurring controls.
FAQ
How often should we run a backup restore test?
Quarterly is a common cadence for small companies, and it is the default in Compliance in a Box. Your auditor decides what is sufficient for your report.
Is a DR test required if we don't include Availability?
A1.3, which covers recovery testing, is only in your report if you choose Availability. CC9.1 still asks you to manage business disruption risk, and many auditors like to see the BCDR plan exercised at least as a tabletop. Ask yours early.
What if the restore test fails?
Record it honestly, fix the cause, and test again. A failed test with a documented fix is better evidence than a quarter with no test at all.