When 8.5 million Windows devices went offline on 19 July 2024 after CrowdStrike pushed a faulty channel file update, organisations running a single-region architecture with no failover path discovered their DR strategy was more theory than practice. The incident cost affected businesses an estimated $5.4 billion in direct losses, according to Parametrix Insurance reporting cited by Reuters. Picking the wrong strategy, or assuming any strategy is inherently resilient, is expensive.
This guide explains the four core disaster recovery strategies, what each one actually costs in architecture and time, and how to match a strategy to your organisation's real recovery requirements.
What is a disaster recovery strategy?
A disaster recovery strategy is a defined approach for restoring IT systems, data, and services after a disruptive event. It specifies where recovery infrastructure lives, how current the backup data is, and how quickly operations can resume. Strategies range from periodic backups with manual restoration to fully mirrored active-active environments operating in parallel at all times.
The four core strategies, compared
Most practitioners and cloud providers organise DR options along a spectrum of cost versus recovery speed. The framework below maps to the AWS disaster recovery whitepaper and is broadly consistent with NIST SP 800-34 guidance on contingency planning for federal information systems.
Before comparing strategies, two numbers matter most: your Recovery Time Objective (RTO) and Recovery Point Objective (RPO). If those terms need grounding, the RTO vs RPO explainer covers the definitions and how to set realistic targets.
| Strategy | Typical RTO | Typical RPO | Relative cost | Best suited for |
|---|---|---|---|---|
| Backup and restore | Hours to days | Hours to days | Low | Non-critical workloads, long recovery windows acceptable |
| Pilot light | 30 min – 4 hours | Minutes to hours | Low-medium | Core systems with infrequent failure risk |
| Warm standby | 5 – 30 minutes | Seconds to minutes | Medium-high | Business-critical applications |
| Active-active (multi-site) | Near-zero | Near-zero | High | Mission-critical, zero-tolerance downtime |
Backup and restore
The most common strategy by deployment count. Data is backed up on a schedule, hourly, daily, or weekly, and systems are rebuilt from scratch when disaster strikes. Nothing runs in the recovery environment until it is needed.
The appeal is cost. You pay for storage, not running compute. The problem is time. Rebuilding from backup typically takes hours, and the February 2024 Change Healthcare ransomware attack illustrated what happens when restore times collide with critical dependencies: UnitedHealth's subsidiary took weeks to restore systems, leaving pharmacies unable to process prescriptions and exposing the strategy's limits at scale.
Backup and restore suits development environments, archival workloads, and any system where an RTO measured in days is genuinely acceptable to the business.
Pilot light
A minimal version of the core system runs continuously in the recovery environment, typically just the database replication and a few critical services. When disaster hits, you scale out the surrounding compute around the pre-warmed core.
The name comes from gas heating: a small flame burns constantly so the boiler can ignite quickly. RTOs in the 30-minute to four-hour range are achievable, but that window assumes scripted automation. Manual scale-out in a crisis adds time and error.
Pilot light works well for organisations that have identified a handful of tier-1 systems but cannot justify full warm standby costs across the entire estate. A clear business impact analysis is a prerequisite, you need to know which services belong in the pilot before you can design it.
Warm standby
A scaled-down but fully functional replica of the production environment runs continuously. At failover, the replica scales up to handle full production load. Unlike pilot light, all application layers are already running; the question is capacity, not presence.
The cost gap between warm standby and pilot light is real. Running even 20-30% of production capacity 24/7 adds up. But for financial services, healthcare, or any sector with regulatory obligations around availability, the trade-off is usually justified.
Under the EU's Digital Operational Resilience Act (DORA), financial entities must demonstrate that critical ICT services can be restored within defined timeframes during disruptions. Warm standby is a practical architecture for meeting those requirements. See the DORA compliance guide for what regulators expect in practice.
Active-active (multi-site)
Two or more fully operational environments handle live traffic simultaneously. Failover is automatic and near-instantaneous because there is no secondary environment waiting, both sites are primary.
This is the most expensive architecture to build and operate. It demands synchronous data replication, global load balancing, and careful attention to data consistency across sites. But for payment processors, trading platforms, or any service where seconds of downtime translate directly to financial or reputational loss, active-active is the only strategy that delivers genuine resilience.
The Snowflake customer breaches between April and June 2024, where attackers used stolen credentials to exfiltrate data from tenants without multi-factor authentication, serve as a reminder that architectural resilience and security controls are separate concerns. An active-active setup does not protect you from credential compromise; it only addresses availability. Reviewed here by Wired.
How to choose: a practical decision path
Start with the business, not the technology. Pull your RTO and RPO targets from your business impact analysis or from contractual SLAs. Then work backwards.
If your RTO exceeds 24 hours, backup and restore is likely sufficient. Invest the budget difference in testing, not infrastructure. An untested backup is not a DR strategy.
If your RTO is 4-24 hours, pilot light with automation is the usual fit. Automate the scale-out scripts, document them, and test them regularly.
If your RTO is under 4 hours, warm standby is the practical floor. Anything with hard regulatory deadlines sits here or above.
If your RTO approaches zero, active-active is required. Budget accordingly and factor in the operational complexity of running synchronised multi-site infrastructure.
Two questions that often get skipped: How frequently will you test this strategy? And does your team know how to execute the failover under pressure, not just in theory? A warm standby that has never been failed over is closer to a pilot light in practice. The DR testing best practices article covers frequency and test types in detail.
Cloud-native considerations
Cloud DR has changed the economics of warm standby and pilot light dramatically. Pay-as-you-go compute means you can spin up recovery capacity in minutes without pre-provisioning hardware. But it introduces a different risk: runaway costs from misconfigured automation, or failover paths that have never been validated because "the cloud handles it."
Cloud DR options are explored in depth in the cloud disaster recovery guide. For multi-cloud or hybrid environments, the critical dependencies mapping process is worth running before committing to an architecture.
Common mistakes to avoid
The most frequent gap practitioners see is the mismatch between documented strategy and actual capability. An organisation documents warm standby, runs backup and restore in practice, and discovers the difference during an incident.
Second: treating DR strategy as a one-time decision. System architectures change. Acquisitions add complexity. A strategy calibrated for 2021 infrastructure may be inadequate for 2024 dependencies. Build in an annual review tied to your DR planning cycle.
Third: underestimating the human side. When DP World Australia suffered a cyberattack in November 2023 and halted port operations for five days, the technical recovery was complicated by unclear escalation paths and unfamiliar runbooks. Technology restores systems; people execute the process. Tabletop exercises and simulations close that gap before an incident does.
Aligning DR strategy with broader resilience
Disaster recovery strategy does not exist in isolation. It sits inside a broader business continuity management framework that accounts for people, facilities, suppliers, and communication, not just IT systems. The BCM program guide covers how these layers interact.
For organisations dealing with cascading failures, where one outage triggers downstream failures across dependent systems, the cascading crises piece offers a practitioner perspective on why single-system DR plans break down in complex incidents.
Choosing a DR strategy is a business decision dressed up as a technical one. Get the RTO and RPO right, test the strategy under realistic conditions, and revisit it when your architecture or risk profile changes.




