Ask most business owners whether they have backups and the answer is yes. Ask how much data they could lose, how long recovery would take, and how far back they can go, and the answer is usually a pause. Those three questions are the heart of a backup policy, and they have names: recovery point objective, recovery time objective, and retention.
You do not need to be technical to set them. You need to understand your business well enough to answer a few questions honestly.
Recovery point objective: how much data can we lose?
The recovery point objective, or RPO, is the maximum amount of data loss you can tolerate, measured in time. If your RPO is 24 hours, you accept losing up to a day of work. If a server fails at 4 p.m. and the last backup ran at midnight, you lose 16 hours of data.
How to choose it
Think system by system:
- Accounting and job cost. Re-entering a day of invoices, timecards, and payments is painful and error prone. Many firms want an RPO of a few hours or less.
- Email. Losing even an hour of messages can mean missing a customer's request.
- Project files. Consider how much work an estimator or project engineer produces in a day.
- Equipment and controller configurations. These change rarely, so the RPO may be days or weeks, but each change should trigger a new backup.
A shorter RPO means more frequent backups and usually a higher cost.
Recovery time objective: how long can we be down?
The recovery time objective, or RTO, is how long a system can be unavailable before the damage becomes unacceptable. It includes everything: deciding to restore, finding the backup, rebuilding, restoring, and testing.
Consider the cost of waiting. If payroll is due Friday, accounting cannot be down for a week. If a pipeline monitoring system is offline, you may need manual operation within minutes. If an archived project folder is unavailable for two days, perhaps nobody notices.
Be realistic
RTO is not what you wish for. It is what your backups and processes can actually deliver. Test it. If a full server restore takes 14 hours when you assumed four, that is useful information.
Retention: how far back can we go?
Retention is how long you keep backups. It matters because problems are not always discovered immediately. Ransomware may sit quietly for days before activating. A file may be quietly corrupted or deleted weeks before anyone looks for it.
Typical considerations:
- Short-term, frequent backups for recent mistakes.
- Weekly and monthly backups kept longer, for problems discovered late.
- Longer retention for records you must keep, such as contracts, safety records, tax and payroll records, and project documentation. Ask your attorney and accountant what applies to you.
Retention costs storage, so match it to the value of the data.
Putting it into a one-page policy
For each major system, record:
- What it is and who owns it.
- RPO and RTO, agreed with the business owner of the system.
- Backup frequency and method.
- Where copies are stored, including at least one offsite and one offline or immutable copy.
- Retention schedule.
- Who monitors backup success and how failures are reported.
- Restore test schedule and results.
- Who is authorized to restore, and who must approve.
The 3-2-1 idea, updated
A long-standing rule of thumb is three copies of your data, on two different types of storage, with one stored offsite. Ransomware has added a modern twist: at least one copy should be offline or immutable, so that stolen administrator credentials cannot delete it.
Common mistakes
- Never testing restores. The most common and most expensive.
- Backups on the same network, with the same credentials, that ransomware can reach.
- Ignoring cloud data. Microsoft 365 and other software-as-a-service tools need their own protection plan.
- Forgetting field and edge systems, such as controller programs, camera recorders, and jobsite servers.
- No alerts, so failed backups go unnoticed for weeks.
- No documentation, leaving recovery dependent on one person's memory.
Match tiers to value
You do not need the same policy for everything. A simple tiering helps:
- Tier 1: Critical systems, short RPO and RTO, tested quarterly.
- Tier 2: Important systems, daily backups, tested twice a year.
- Tier 3: Lower priority or archival, less frequent backups, tested annually.
Review annually
Business changes. New projects, new software, and new locations alter what must be protected. Review the policy each year and after any significant change or incident.
Working with Ironfield Cyber
Ironfield Cyber helps contractors and energy companies set recovery objectives with their leadership team, design backups to meet them, and test the results. If you are not sure what your current RTO really is, a timed restore test will tell you.