Skip to main content
IT KORR
IT KORRKeeping Organizations Reliable & Resilient
Guidance

Disaster Recovery Testing

Disaster recovery testing is the practice of deliberately exercising a recovery capability to confirm it works, rather than assuming it does. It runs across five levels, from verifying that a backup produces usable data through to operating the business from the recovery position — and the distinction that matters most is between testing a backup and testing a recovery.

The first verifies that data exists and can be read. The second verifies that the organisation can resume operating — which involves dependencies, sequence, identity, network, people and elapsed time. Almost every organisation does the first and records it as if it were the second.

Test Types

Five levels, and what each one actually proves

These are cumulative rather than alternatives. An organisation that has never completed a component recovery test should not be planning a full failover.

  • 01

    Restore verification

    What it proves
    That a specific backup contains recoverable data.
    Effort and disruption
    Low — minutes to hours, no disruption
    Worth running
    Monthly, sampled across systems

    Restore a defined file, mailbox item or database to an isolated location and confirm it opens, parses and contains what it should. This is the floor rather than a real DR test, and it is the only one on this list most organisations actually perform. It catches silent corruption and misconfigured job scope, and it proves nothing about recovery time.

  • 02

    Tabletop exercise

    What it proves
    That people know the plan and that the plan matches the environment.
    Effort and disruption
    Low — two to three hours, no technical disruption
    Worth running
    Annually, and after any material change

    Walk a scenario through with the people who would actually respond. The value is almost never in the technical answers; it is in discovering that the plan names a system retired two years ago, that the person listed has left, or that two people each believed the other would make the call. Cheapest test on the list and the one that most reliably finds something.

  • 03

    Component recovery test

    What it proves
    That one real system can be recovered to a working state.
    Effort and disruption
    Medium — a scheduled window, limited disruption
    Worth running
    Quarterly to annually, rotating systems

    Recover one production system into an isolated environment and bring it to a usable state — not just restored, but started, authenticated against, and checked. This is where dependency surprises surface: the application restores and cannot reach a domain controller, a licence server, or a database nobody listed as a dependency.

  • 04

    Parallel / isolated failover

    What it proves
    That a group of dependent systems can run together away from production.
    Effort and disruption
    High — significant planning, isolated network required
    Worth running
    Annually where recovery objectives are aggressive

    Stand the recovery environment up alongside production, isolated so it cannot interfere, and exercise the workflows that matter end to end. This is the first test that genuinely measures recovery time rather than estimating it, and the measured number is routinely several times the assumed one.

  • 05

    Full failover

    What it proves
    That the organisation can actually operate from the recovery position.
    Effort and disruption
    Very high — real disruption, real risk, executive sign-off
    Worth running
    Rarely; where regulation or contract requires it

    Move production to the recovery environment and run the business from it. The only test that proves the whole chain including the human and process parts. It carries genuine risk and should be approached only once the lower tiers pass consistently — an organisation that has never completed a component test should not be attempting this.

Evidence

Six things a test record has to contain

A record that cannot demonstrate these is not evidence of testing, whatever it is titled. This is also, almost exactly, what an insurer or auditor asks to see.

  • What was tested, specifically

    A record saying "DR test performed" proves nothing. Name the system, the data set, and the recovery point used.

  • Who performed it

    Ideally whoever would perform it in a real incident. A test run by the one engineer who built the system measures that engineer, not the organisation.

  • Elapsed time, measured not estimated

    This is the number that determines whether the stated RTO is real. It is also the number most often absent from test records, because measuring it is inconvenient.

  • The result against a stated objective

    A test with no pass condition cannot fail, which is why written RPO and RTO per workload have to exist before testing is meaningful.

  • What went wrong, and what changed

    The most valuable line in the record, and the most commonly omitted. A test that found nothing is usually a test that was scoped to find nothing — and a reviewer reads a clean record with more suspicion than a messy one.

  • A date

    Obvious, routinely missing. An undated record cannot demonstrate cadence, and cadence is what insurers and auditors are assessing.

Failure Modes

Six ways a disaster recovery test proves nothing

Each of these produces a test record. None of them produces confidence.

  • Testing the backup instead of the recovery

    Confirming a job completed, or that a file restores, and recording it as a DR test. It verifies the backup and says nothing about whether the organisation can resume operating. These are different questions and the second is the one being asked.

  • Testing the system that is easiest to test

    The file server gets tested every year; the line-of-business application with eleven dependencies never does. Test coverage drifts toward whatever is convenient, which is precisely the inverse of where the risk sits.

  • Identity left out of scope

    The single most common dependency failure. Systems restore and nobody can authenticate, because directory services, MFA infrastructure or Conditional Access were never considered part of the recovery scope. A recovery plan that assumes identity is available has assumed away the hardest part of a ransomware recovery.

  • The recovery environment shares a failure domain with production

    Backups reachable with production administrative credentials, or a recovery site dependent on the same directory, network path or cloud tenant. The test passes because the shared dependency was available, and the real event is exactly when it is not.

  • No stated objective, so no pass condition

    Testing without written RPO and RTO per workload produces a result nobody can interpret. Four hours might be excellent or catastrophic; without a target the test generates activity rather than information.

  • SaaS assumed to be out of scope

    Microsoft 365 holds the operational data of most organisations and is routinely absent from DR planning entirely, on the assumption that the platform handles it. The platform provides retention, not backup, and its default windows are measured in days.

FAQ

Common Questions

What is disaster recovery testing?

Disaster recovery testing is the practice of deliberately exercising an organisation’s recovery capability to confirm it works, rather than assuming it does. It spans five levels: restore verification (can this backup produce usable data), tabletop exercise (do people know the plan and does it match reality), component recovery (can one real system be brought back to a working state), parallel failover (can dependent systems run together away from production), and full failover (can the business actually operate from the recovery position). The distinction that matters most is between testing a backup and testing a recovery — the first verifies data exists, the second verifies the organisation can resume operating.

How often should disaster recovery be tested?

By test type rather than as a single cadence. Restore verification monthly, sampled across systems; a tabletop exercise annually and after any material change to the environment or the team; component recovery tests quarterly to annually on a rotation so coverage does not drift to the easy systems; parallel failover annually where recovery objectives are aggressive. Full failover is rare and generally driven by regulation or contract. The more useful trigger than the calendar is change: a significant infrastructure, identity or provider change invalidates the assumptions the last test confirmed.

What is the difference between a backup test and a disaster recovery test?

A backup test asks whether the data is there and restorable. A disaster recovery test asks whether the organisation can resume operating — which involves dependencies, sequence, identity, network, people and elapsed time. Backup jobs reporting success is the single most common thing mistaken for recovery readiness, and the gap between them is only ever discovered at the worst possible moment.

What evidence does a disaster recovery test need to produce?

Six things: what specifically was tested (named system, data set, recovery point), who performed it, measured elapsed time rather than an estimate, the result against a stated recovery objective, what went wrong and what changed as a result, and a date. The "what went wrong" line is the most valuable and the most commonly omitted — an experienced reviewer reads a perfectly clean test record with more suspicion than a messy one, because real tests find things.

Does cyber insurance require disaster recovery testing?

Tested recovery has become a routine condition of placement and renewal rather than a differentiator, and underwriters increasingly ask for evidence rather than an attestation. In practice what gets requested is a dated restore-test record and stated recovery objectives. This is one reason DR testing has moved from an IT hygiene topic to a commercial one — confirm the specific requirements with your broker, since they vary by carrier and by policy.

Should Microsoft 365 be included in disaster recovery testing?

Yes, and it usually is not. Microsoft 365 holds the operational data of most organisations and is routinely excluded from DR planning on the assumption that the platform handles recovery. It provides retention, not backup: 14 days by default for purged Exchange items and 93 days across both recycle-bin stages for SharePoint and OneDrive. Any DR test that excludes the platform holding the organisation’s email and documents is testing the smaller half of the problem.

What is the most common disaster recovery test failure?

Identity. Systems restore correctly and nobody can authenticate, because directory services, MFA infrastructure or Conditional Access were never scoped into the recovery. It is especially consequential in ransomware scenarios, where identity infrastructure is frequently the thing compromised — so a recovery plan that assumes identity is available has assumed away the hardest part of the actual event.

Can disaster recovery testing be done without disrupting production?

The lower three tiers, yes. Restore verification and tabletop exercises carry effectively no production risk, and component recovery into an isolated environment carries very little when the isolation is genuinely enforced. Parallel failover needs careful network isolation but does not need to touch production. Only full failover is inherently disruptive, which is why it should be reached rather than started with.

Recovery Readiness

Find out what your environment would actually restore

A review of backup coverage against a real asset inventory, retention against your actual obligations, isolation from the production administrative boundary, and whether a restore has ever been performed end to end. You get the findings either way.

We respond within one business day.

Build: 4fb1bc8 | Built: Oct 6, 2026 8:27 PM EDT