Skip to content
Backup & continuity

Testing your backups: the method to be sure you can restore

Backup restore testing for businesses: file, VM, application, full site. Test calendar, real time vs RTO, test reports, ransomware scenario.

By the ALLSAFE SOLUTIONS engineering team30 September 20267 min read
Testing your backups: the method to be sure you can restore

Every morning, the backup report shows green lines. Yet in the audits we run, the question “when did you last restore a server?” often goes unanswered. The short answer: a backup is only reliable once a restore has actually been tested, timed and documented. That takes four levels of testing, a calendar set in advance, a comparison of real restore time with your RTO and a simple rule for when a test fails. Here is the method we apply in the companies we support in Morocco.

Why “backup succeeded” does not mean “restorable”

Backup software reports that it copied blocks without error. By default, it does not check that those blocks form a system that boots and an application that works. The causes of failure we see at restore time are rarely spectacular:

  • an inconsistent database, backed up mid-write, without application-aware processing;
  • a forgotten volume: the data disk was added after the job was created and is not in it;
  • a missing dependency: the application server comes back, but the directory or licence server was not restored;
  • an encryption key nobody can find: the backup is protected by a password no one remembers;
  • a restore far slower than expected, because the internet link or standby storage was never measured under real conditions.

None of these shows up in the daily report. Only a test reveals them. This is the “0” in the 3-2-1-1-0 rule: zero errors found during a real restore.

The four levels of testing

Testing backups is not just recovering a file. Each level proves something different, and none replaces the others.

Level What you do What it proves What it does not prove
File Restore a specific document or folder to a given date Data is readable and retention works That a full server can restart
Virtual machine Boot a VM from the backup in an isolated network The system boots, disks are complete That the application and its data are consistent
Application Restore the ERP, email or database and have a key user check the data The business can actually resume on that service That everything fits within the target time
Full site Rebuild everything in the DRP order, stopwatch in hand The real time back to normal Nothing more: this is the reference test

The file test is useful but misleading when it is the only one performed: it reassures without proving you can bring a server back. The full-site test, on the other hand, is heavy; it is not run every week, but it is the only one that measures your real recovery capability.

Automate what you can: SureBackup and Instant Recovery

With Veeam, part of the testing can be automated. SureBackup starts backed-up machines in a virtual lab isolated from production, then checks that they respond: system boot, network response and, if configured, application-specific test scripts. A scheduled job thus flags an unusable backup well before the day you need it.

Instant Recovery lets you start a virtual machine directly from the backup file while the final restore continues in the background. It is a fast recovery tool, but also an excellent test: if a critical VM will not start in Instant Recovery, you will know before the incident.

These automations do not replace human judgement. A lab confirming that a server “responds” does not tell you whether last month’s accounting is complete. That is the role of the application test, carried out with a business user. The Veeam settings that make these tests possible are covered in our guide Veeam: the 6 settings.

A test calendar for your business

The right pace depends on how critical each system is. Here is the framework we suggest to SMBs, adjusted during the audit:

  • Continuously: read backup alerts and follow up as soon as a job fails or stays in warning.
  • Every week: automated check of critical machines in an isolated lab.
  • Every month: restore randomly chosen files, including in Microsoft 365 (an email, a OneDrive file, a SharePoint site).
  • Every quarter: full application restore of at least one critical system, timed and validated by a key user.
  • Every year: disaster exercise across the whole scope, in the order set by the recovery plan.
  • After every major change: new server, migration, change of storage or provider.

Rotate the systems tested from one quarter to the next to cover the whole scope within the year.

What to measure: real time against RTO

A test without a stopwatch is only half useful. The disaster recovery plan sets two targets: the RTO, acceptable downtime, and the RPO, acceptable data loss. The test gives their real values.

Measure How to get it Question to ask
Real restore time Stopwatch, from launch to validation by the user Is it below the RTO you set?
Real RPO Date and time of the restored point compared with the test time Would the data loss be acceptable?
Completeness Check by a key user: latest entries, attachments, access rights Is any data or dependency missing?
Manual steps List of undocumented actions improvised during the test Would someone else have known how to do them?

Measure end to end, not just copy time: finding the right restore point, access, network reconfiguration and validation are all part of real downtime.

Document every test

An unrecorded test never happened, for an auditor as for your successor. Each report fits on one page: date, system tested, restore point used, who restored, who validated, time measured, result, gaps found and actions decided. Keep these reports together: they show how your recovery capability evolves and serve as evidence during an inspection, for example on personal data protection.

When a test fails

A failed test is good news: the problem was found before the incident. The rule is simple: until the test succeeds, consider that backup as missing.

  1. Open an incident, with the same priority as an outage.
  2. Identify the cause: configuration, storage, credentials, procedure or dependency.
  3. Fix it, then rerun the same test until it succeeds.
  4. Check whether the same cause affects other jobs.
  5. Update the restore procedure and the recovery plan.

The ransomware scenario

The most demanding test simulates ransomware. It adds three constraints: the production network and directory are assumed compromised, the restore starts from the immutable or off-site copy, and the chosen point must predate the intrusion. So check that you can reach the backups without domain accounts, that restore time from the remote copy fits your RTO, and that restored data is scanned before going back into service, so the attacker is not reintroduced.

The most common mistakes

  • Always testing the same small file, which proves nothing about servers.
  • Testing in production, risking overwriting real data or causing address conflicts.
  • Never timing the restore, and discovering during the incident that it exceeds the RTO.
  • Forgetting dependencies: directory, DNS, licences, certificates.
  • Having a single person test, who becomes indispensable on disaster day.
  • Not validating with the business: the machine boots, but nobody checked the data.

Checklist: your restore tests

  • All four test levels are covered: file, VM, application, full site.
  • A written calendar states who tests what, and when.
  • Tests run in a network isolated from production.
  • Every test is timed and compared with the RTO and RPO.
  • A business user validates the restored data.
  • A one-page report is archived for each test.
  • Every failure is handled as an incident and the test rerun.
  • A ransomware scenario from the immutable copy has been run at least once.
  • At least two people know how to restore critical systems.

How we do it

We start with an inventory: backup jobs, scope covered, date and result of the last real test. We then define with you the test calendar and RTO and RPO targets per system, and set up automated checks and the isolated lab. Quarterly restores are carried out with your key users and recorded in the service report. Alerts feed into our 24/7 monitoring, and a critical incident is handled in under 15 minutes. This support is part of our backup and business continuity offering; our case studies show comparable projects.

Not sure your backups can be restored? The initial audit is free, and we reply within 24 business hours: contact us.

Frequently asked questions

Why can a backup marked "successful" be impossible to restore?

The status says the copy finished, not that the data is usable. An inconsistent database, a missing boot disk, a lost encryption password or a forgotten dependency such as the directory only show up when you try to restore.

How often should you test restoring your backups?

A frequent automated check, for example weekly with a tool like Veeam SureBackup, an application restore every month or quarter depending on criticality, and a timed full restore at least once a quarter, plus a yearly disaster exercise.

How do you test a restore without disrupting production?

By restoring into an isolated network with no link to production: machines start with their original addresses without conflict, and key users check the data on a copy. Nothing is written to the real systems.

What should you do if a restore test fails?

Treat the failure as an incident: find the cause, fix the configuration or procedure, rerun the same test until it succeeds, then update the disaster recovery plan. Until the test succeeds, consider that backup as missing.

About the editorial team

ALLSAFE SOLUTIONS

Network, security and cloud engineers

Written by the engineering team at ALLSAFE SOLUTIONS, a managed IT provider founded in Casablanca by network, security and cloud engineers. Our articles draw on the projects we deliver for clients in Morocco and abroad.

LinkedIn
← All articles
CallWhatsAppFree audit