How many backups do you actually need? A retention strategy guide
Quick answer
Start from the industry-standard 3-2-1 rule: 3 copies of your data, on 2 different types of media or storage, with 1 copy offsite. Then set retention windows based on how far back you'd realistically need to restore from: days for accidental deletion, weeks for slow-discovered corruption, longer for compliance-driven needs. And treat an untested backup as an assumption, not a guarantee, until you've actually restored from it once.
The 3-2-1 rule as a starting point
The 3-2-1 rule is a long-standing, industry-standard baseline for backup strategy, not specific to any one product or provider: keep 3 copies of your data in total, store them on 2 different types of media or storage systems, and keep 1 copy offsite, away from the primary location. The logic behind each number is straightforward. Three copies means the original plus two backups, so a single failure never leaves you with zero redundancy. Two different media types protects against a failure mode that could take out an entire class of storage at once, a RAID controller fault or a specific storage system bug, for instance. One copy offsite protects against anything that affects the whole primary site: fire, theft, a datacenter-level incident, or simply an outage that takes the primary location offline entirely.
3-2-1 is a starting framework, not a rigid formula. What it forces you to think through is where a single point of failure could still take out both your live data and your backups at once, and to deliberately break that link.
Setting retention windows: how far back would you actually need to go?
Retention isn't one number, it's really several different problems that happen to use the same word. Work out, separately, how far back you'd need to restore from for each of these situations, because they call for different retention windows:
- Accidental deletion. Someone deletes the wrong file or drops the wrong table. This is usually noticed quickly, so a retention window measured in days is often enough, as long as backups are frequent enough that you're not losing much work between the deletion and the most recent backup before it.
- Slow-discovered corruption. Some problems, a subtly corrupted database, a bad migration, a bug quietly writing wrong data, aren't noticed for a while. By the time you spot it, your most recent backups may already contain the same corruption. This calls for a retention window measured in weeks, so you have a clean point to restore from that predates when the problem started.
- Compliance-driven retention. Some data has to be retrievable for a much longer period for regulatory or contractual reasons, independent of any actual operational incident. This is a different requirement again, usually months or longer, and it's driven by policy rather than by how quickly problems tend to surface.
Treating these as one setting usually means either paying to retain everything for the longest window unnecessarily, or discovering too late that the specific restore point you needed already aged out.
Snapshot-style backups vs. capacity-oriented backup storage
These solve different parts of the same problem, and most real strategies use both rather than picking one:
- Snapshot-style backups are fast and frequent, and usually kept for a short retention window. They're built for quick rollback: undo the last few hours or days of change with minimal fuss. See Backups and snapshots: what's the difference for how this works in practice.
- Capacity-oriented backup storage is less frequent and built for longer retention, and it's the more cost-efficient place to keep data you need to hold onto for weeks, months, or longer, rather than for a same-day rollback. See Using Backup Storage as a disaster-recovery target.
Matching the right tool to the retention problem it's actually solving, quick rollback versus long-window recovery, keeps both cost and complexity down compared with trying to make one system do both jobs well.
Why testing a restore matters as much as taking the backup
A backup you've never restored from is an assumption, not a guarantee. Backups can fail silently: a job that reports success but backed up an already-corrupted file, permissions that block a restore when you actually need one, an incomplete backup that nobody noticed because nobody looked. The only way to know a backup is actually usable is to restore from it and check the result, and doing that occasionally, not just when a real incident forces the question, is what turns "we have backups" into something you can actually rely on under pressure. This matters just as much when you're winding a service down: see What happens to your data when you cancel a service for what you need to have already secured elsewhere before that point.