Most companies that experience a serious outage discover something about their backups they did not know: a job had been failing for weeks, a key database was never included, or the restore takes far longer than anyone assumed. The way to find these problems on your own schedule instead of during a crisis is to run restore drills.
A drill does not need to be dramatic or expensive. It needs to be regular, written down, and honest.
What a restore drill proves
A drill answers questions that a green status light cannot:
- Is the data actually in the backup?
- Can we restore it to a working state?
- How long does it take?
- Do the right people know what to do?
- Do we have the credentials, licenses, and documents needed?
Three levels of testing
Level 1: File restore
Choose a few files or a folder and restore them to a different location. Open them and confirm they are complete. This is quick and should be done monthly.
Level 2: System restore
Restore a full server or database to a test environment. For an accounting system, confirm that the application starts and that the data looks current. Do this quarterly for your most important systems and at least annually for the rest.
Level 3: Scenario exercise
Simulate a major event, such as losing the main server or the entire office, and walk through the recovery plan, including communication and decision making. Run it annually. Involve operations and finance, not only IT.
A simple drill, step by step
- Pick the target. Choose a system, such as the accounting database or file server, and choose a restore point from a few weeks ago.
- Assign roles. Someone performs the restore, someone times it, and someone from the business verifies the result.
- Set the clock. Record when you start.
- Restore to an isolated location. Never overwrite production data in a test.
- Verify. Business users check that records, attachments, and recent transactions are present and usable.
- Record the time and issues. Note how long it took and what went wrong.
- Compare to your recovery objectives. If the business needs to be running in four hours and the restore took two days, you have found a gap worth fixing.
- Fix and repeat. Update procedures, correct the backup configuration, and schedule the next drill.
What to check beyond the data
Access to the backup itself
Who can log into the backup system? Are those credentials stored somewhere you can reach if the network is down? Is multi-factor authentication in place, and do you have a way around it in an emergency?
Encryption keys and licenses
If backups are encrypted, losing the key means losing the backup. Store keys securely in more than one place. Confirm software licenses allow restoring to new hardware.
Dependencies
Many systems depend on others, such as identity services, DNS, or a database server. Test restoring them in the correct order.
Documentation
Keep printed or offline copies of the recovery runbook, including network diagrams, vendor contacts, and account details. Your documentation system may be unavailable during an incident.
Off-site and immutable copies
Test restoring from the copy an attacker could not reach, not only the convenient local one. That is the one you will rely on in the worst case.
Common failures drills reveal
- Backup jobs that silently fail or skip open files.
- Systems added after the backup was set up and never included.
- Retention too short to reach a clean point before an infection.
- Restore times far longer than expected because of bandwidth limits.
- Only one person knows how to do the restore.
Make it routine
Put the drills on the calendar, assign owners, and report results to leadership in plain language: what was tested, how long it took, what was fixed. Treat a failed drill as a success of the process, since it found the problem early.
Working with Ironfield Cyber
Ironfield Cyber monitors backups and runs scheduled restore tests for contractors and energy companies, with reports a non-technical owner can read. If you cannot recall the last time you restored anything, we can help you run your first drill.