Backup Testing: 7 Critical Failures a Restore Drill Exposes

Backup testing gap chart: 90% of organisations were confident of recovery but only 28% recovered all their data after a ransomware attack
Confidence is measured before the incident. Everything else is measured after.

Almost every business we assess says it has backups. Far fewer can name the date somebody last restored one and watched it come back. That gap — between a job that reports success and a system that actually returns — is what backup testing exists to close, and it is the control most commonly missing from a small estate.

This is what a real restore drill involves, what it reliably exposes, and how often it is worth doing.

What backup testing actually is

A backup job reporting success tells you that data was written somewhere. It does not tell you the data can be read, that it is complete, that the application it belongs to will start, or how long any of that takes.

Backup testing is the exercise that answers those questions before an incident does. It has a scenario, a stopwatch, a definition of success agreed in advance, and a written result.

The distinction matters because the two are usually confused. Monitoring a backup job is an operational task that runs every day. Testing the restore is a scheduled exercise that a person plans, runs and reports on — and it is the only thing that produces evidence.

The UK’s National Cyber Security Centre puts it plainly in its guidance on offline backups in an online world: backups should not only be created frequently, they should be regularly tested to check they work as expected. The same guidance is blunt about why the offline copy matters, having watched ransomware take out connected USB drives, network shares and cloud sync targets alongside the originals.

The 2026 numbers: confidence is running ahead of capability

Veeam surveyed more than 900 senior IT, security and risk leaders for its Data Trust and Resilience Report 2026, published in April. Ninety per cent said they were confident they could recover from a cyber incident within their recovery time objectives.

Then the same report gives the outcome. Organisations that suffered a ransomware attack recovered, on average, 72% of the affected data. Forty-four per cent recovered less than three quarters of it. Only 28% recovered all of it.

Two thirds of the confidence is not supported by the results. That is a measurement problem before it is a technology problem, and backup testing is the measurement.

Seven failures backup testing reliably exposes

These are the ones that come up repeatedly. None of them is exotic, and none of them is visible from a green dashboard.

1. There is nothing to restore onto

The British Library’s published review of its 2023 ransomware attack is the clearest public example there is. The Library had secure copies of its digital collections and metadata. Its own words: “we have been hampered by the lack of viable infrastructure on which to restore it”.

The attackers had destroyed servers specifically to inhibit recovery. Major systems could not be brought back in their pre-attack form at all, because vendors no longer supported them or because they would not run on the new, secured infrastructure. Having the data was never the problem — and only backup testing finds that out in advance.

2. The backup is a sync, and a sync replicates the damage

File sync services faithfully propagate encryption, deletion and corruption. That is their job. They are not a second copy in any sense that survives an incident.

Cloud platform retention is also shorter than people assume. Microsoft’s own documentation puts the default deleted item retention in Exchange Online at 14 days, configurable to a maximum of 30, after which the item is removed. SharePoint and OneDrive hold deleted items for 93 days. Those are recycle bins, not a recovery plan.

3. Retention shorter than the attacker’s head start

Mandiant’s M-Trends 2026 puts the global median dwell time for 2025 at 14 days, up from 11 the year before. For cyber espionage and North Korean IT worker cases the median was 122 days.

If every restore point you hold is newer than the intrusion, every restore point you hold contains the intrusion. Retention has to be measured against how long a compromise plausibly goes unnoticed in your environment, not against how much storage costs.

4. Nobody knows the order

Systems have dependencies. The line-of-business application needs its database, the database needs a domain controller, the domain controller needs DNS, and the whole thing needs an identity provider that may itself be the thing under attack.

Backup testing surfaces that order because it forces you to walk it. A spreadsheet of backup jobs does not.

5. The keys to the recovery are inside the wreck

Backup console credentials in a password manager that is down. The runbook on a file share that is encrypted. Multi-factor prompts routed to a phone system that is offline. Recovery documentation stored in the tenant being recovered.

This one is embarrassing and extremely common, and it costs nothing to fix once seen.

6. The restore works, but not at the speed anyone assumed

Restore rates are governed by throughput, item counts and the type of restore. Microsoft publishes indicative figures for its own Microsoft 365 Backup service — median rates of one to three terabytes per hour, and mailbox restores in the range of 100 to 500 items per minute.

Do that arithmetic against your own data volumes before an incident rather than during one. A recovery time objective chosen in a meeting and never measured is a wish; backup testing is what turns it into a figure.

7. The data is back and the business still cannot work

The last failure is the subtlest. Files restore, services start, and the people who use the system find that the records are stale, the integrations are pointing at the wrong place, or half a day of transactions is missing and nobody knows which half.

The British Library review notes that each restored dataset had to be validated for integrity before it went back. Verification is part of the restore, not an afterthought — and the person qualified to verify it usually does not work in IT.

The three depths of contingency plan test in NIST SP 800-34: tabletop discussion, functional restore from backup media, and full-scale failover and reconstitution
Three depths, three different questions. Most businesses have never done the second.

What good backup testing looks like, step by step

NIST’s SP 800-34 Rev. 1 is the useful reference here, and it is free. It sets out three depths of exercise: a discussion-based tabletop where nothing is touched, a functional exercise that includes system recovery from backup media, and a full-scale exercise that fails over to an alternate location and reconstitutes the system to a known state.

It also lists what a contingency plan test should cover: notification procedures, system recovery on an alternate platform from backup media, internal and external connectivity, system performance on alternate equipment, and restoration of normal operations. Note how little of that is about the backup file itself.

The instruction worth stealing is this: develop a test plan with explicit objectives and success criteria, a scenario that mimics reality as closely as possible, and an after-action report capturing corrective actions. Without success criteria agreed beforehand, every test passes.

A workable backup testing drill for a business without dedicated IT staff looks like this:

  • Choose a scenario, not a file. “The finance server is gone and we do not trust anything in the domain” is a scenario. “Restore a document” is a spot check.
  • Write the success criteria first. Which system, restored to what point in time, working well enough for whom to do what, within how long.
  • Restore somewhere clean. Isolated infrastructure, not over the top of production. Where you cannot afford that, restore a representative subset and say so in the report.
  • Time it honestly. Start the clock when the decision to recover is taken, not when the copy starts. Waiting for approval and hunting for credentials are part of the recovery time.
  • Have the business verify. Someone who uses the system daily confirms the data is right. IT confirming that IT’s restore worked proves very little.
  • Write it down. Date, scenario, what worked, what did not, elapsed time, and the corrective actions with owners.

That last artefact is what a regulator, an insurer or an enterprise client will ask for. It is also what makes the next drill shorter.

How often backup testing should happen

NIST scales the depth to the impact of the system rather than prescribing one cadence for everything: a tabletop is sufficient for low-impact systems, moderate-impact systems get a functional exercise including recovery from backup media, and high-impact systems get a full-scale exercise with failover and full reconstitution.

Insurers have converged on asking the same thing as a frequency rather than a yes or no. Beazley’s published cyber application form offers never, annually, two to three times a year, or quarterly or more often. There is no box for “we assume it works”. We walk the rest of that form question by question in our piece on the 12 controls underwriters check.

Our own view of a sensible backup testing cadence, for a business of twenty to two hundred people: a file-level restore every month as a spot check, a full functional restore of one important system every quarter, and one scenario-driven exercise a year that includes the people who would have to make decisions. Rotate which system gets the quarterly drill so that everything is exercised inside a year.

When backup testing is the wrong thing to do next

The unprofitable thing to say, so we will say it.

If you have no offline or immutable copy at all, stop and fix that first. Testing a backup that a ransomware operator can reach and delete tells you only that it works today, under conditions that will not apply on the day it matters.

If your critical application is unsupported and would not run on rebuilt infrastructure, backup testing will produce a report you cannot act on. That is a lifecycle problem, and the British Library’s own lesson list puts it exactly there: legacy systems are not merely hard to secure, they are extremely hard to restore.

And full-interruption testing is genuinely dangerous for a small business with no spare capacity. NIST wants disruptive tests approved by the senior person accountable for the system, and that is the right instinct. A drill that takes out a working environment has converted a hypothetical outage into a real one.

We are not going to publish a cost-per-hour of downtime, and you should be wary of anyone who does. The per-minute figures that circulate in vendor material are drawn from surveys of very large enterprises and do not transfer to a fifty-person business. Work out your own: hourly payroll of the people who would be idle, plus the revenue that does not get invoiced, plus the contractual penalties, plus the cost of the recovery itself. It takes an afternoon and the number will be yours.

Backup testing, POPIA and the client questionnaire

South African businesses have a second reason to care beyond not losing the company.

POPIA’s security safeguards oblige a responsible party to secure the integrity and availability of personal information — availability, not just confidentiality. Data destroyed in an incident you cannot recover from is a failure of that duty, and a restore you have never tested is not evidence of anything. We set out the technical half of that obligation in our POPIA compliance checklist.

The same drill report answers the enterprise security questionnaire, the insurer’s renewal question and the board member who asks how long we would be down. One piece of work, three audiences.

A 90-day backup testing plan

If none of this exists today, this is the order that changes the answer fastest.

Month one. Inventory what is actually backed up, and more importantly what is not — the SaaS platforms, the laptops, the one server somebody set up outside the process. Confirm at least one copy is offline or immutable. Write down a recovery time and recovery point objective per system, even if the first draft is a guess. Note what each backup product talks to while you are there: mail and archiving tools are among the heaviest users of Exchange Web Services, and backup testing is a sensible moment to confirm yours is not caught by the EWS retirement.

Month two. Run one functional restore of one important system to isolated infrastructure. Time it. Expect it to go badly; that is the point, and the first drill is worth more than the next five.

Month three. Fix what the drill found, then hold a tabletop for the decision-makers: who declares an incident, who authorises the spend, who talks to clients, and what the business does manually while systems are down.

After that it is a calendar entry and a short report each quarter. The estate-level version of this conversation — what to fix, in what order, against a budget — is the work we describe under the virtual CIO role.

The five steps of a restore drill: choose a scenario with success criteria, restore to clean infrastructure, time it against the stated RTO, have the business verify the data, write the after-action report
The stopwatch starts when the decision is taken, not when the copy starts.

The summary worth keeping

A backup is a claim. A restore is a fact. Backup testing is how you convert one into the other, on a date of your choosing rather than the attacker’s.

The British Library’s tenth lesson to itself was to prioritise recovery alongside security, on the grounds that no security is perfect and the ability to recover quickly is essential when — not if — an attack succeeds. That is the whole argument, written by an organisation that learned it the expensive way.

If none of this has been tested in your environment, that is the finding, and it is the thing worth fixing first. It is also what our disaster recovery and backup work starts with — recovery objectives agreed with the business, then a drill with the result written down.