What is an operational resilience framework?
An operational resilience framework is a structured set of policies, processes, and controls that enables an organization to identify its most critical business services, define the maximum tolerable level of disruption for each, map the people, technology, data, and third parties that support those services, and continuously test whether the organization can stay within those limits under realistic stress conditions.
The framework does not replace business continuity management. It extends it. BCM provides the foundational artefacts: business impact analyses, recovery plans, and exercise cadences. The resilience framework uses those artefacts to answer a sharper question: can we actually absorb disruption and maintain outcomes for customers and markets?
Why regulators made this formal
After the 2012 RBS payment failure left 6.5 million UK customers locked out of their accounts for days, regulators concluded that recovery plans alone were insufficient. The Bank of England, PRA, and FCA spent the better part of a decade developing what became the UK operational resilience policy, published in March 2021 and requiring full compliance by March 2025.
The EU followed with DORA, the Digital Operational Resilience Act, which became applicable in January 2025 and imposes ICT risk management, incident reporting, and third-party oversight requirements on financial entities across the bloc. Australia's APRA updated CPS 230 in 2023 with mandatory operational risk management standards that took effect in July 2025.
For a deep look at what DORA demands specifically, see DORA compliance: a complete guide for European financial services.
The five core components
Most mature frameworks, whether built around UK FCA expectations, DORA, or ISO 22301, converge on five structural elements.
1. Critical business service identification
This is the starting point. Not every process matters equally. Teams must map the services that, if disrupted, would cause harm to customers, financial markets, or the broader economy. For a retail bank that means payment processing, mortgage origination, and deposit access. For a hospital network it means triage, pharmacy dispensing, and radiology.
The identification exercise draws directly from business impact analysis outputs. BIAs are not optional pre-work; they are the engine that generates the service inventory and the initial prioritization.
2. Impact tolerances
For each critical service, the organization must define a maximum tolerable disruption: how long the service can be unavailable, or how degraded it can become, before harm becomes unacceptable. The UK framework calls these "impact tolerances." DORA uses recovery time objectives and recovery point objectives at the ICT layer, which you can explore further in RTO vs RPO explained.
Setting tolerances is not a technical exercise. It requires input from business owners, legal, compliance, and often regulators themselves. A payment firm might set a four-hour tolerance for its retail payment rails. A trading platform might set fifteen minutes.
3. Mapping people, processes, technology, and third parties
Once services and tolerances are defined, teams map every resource the service depends on: staff, facilities, applications, data, and external vendors. The Change Healthcare ransomware attack in February 2024 illustrated what happens when this mapping is incomplete. UnitedHealth Group's subsidiary processed around 15 billion healthcare transactions annually. When the attack hit, providers across the US lost the ability to submit claims, and dependency chains that nobody had fully documented suddenly became visible through failure.
Third-party mapping deserves particular attention. A single point of failure buried in a vendor relationship can cascade across multiple critical services. The third-party risk dimension of resilience has moved from footnote to regulatory priority.
4. Scenario testing against impact tolerances
Documented plans are necessary but not sufficient. The framework requires testing whether the organization can actually stay within its impact tolerances during realistic adverse scenarios. This is where many programs expose gaps.
Testing approaches range from tabletop exercises to full technology failover tests. Tabletop exercises are the most accessible entry point. Crisis simulations that inject live pressure are more revealing. The BCI's Good Practice Guidelines recommend that scenario design cover both acute shocks and slow-burn stress, not just single-point failures.
The common failure mode: programs measure exercise completion rather than whether the organization stayed within tolerance. For a sharper look at this problem, why most BCM exercise programs measure activity, not readiness is worth reviewing.
5. Governance, reporting, and continuous improvement
A framework without ownership degrades. Boards and senior management must formally own impact tolerances and sign off on the annual self-assessment of resilience. The FCA expects documented board-level engagement, not just delegation to a continuity team.
Reporting structures should connect operational resilience findings to enterprise risk committees. Lessons from exercises, incidents, and near-misses must feed back into service mapping and tolerance reviews. The framework is a cycle, not a project.
How the framework fits with existing BCM disciplines
Practitioners sometimes worry that an operational resilience framework will make their existing BCM work redundant. It does not. It contextualizes it.
| Discipline | Primary question | Key artefacts | Role in the framework |
|---|---|---|---|
| Business continuity management | How do we recover? | BCP, BIA, recovery plans | Foundation layer |
| Disaster recovery | How do we restore systems? | DR plan, RTO/RPO targets | Technology execution layer |
| Crisis management | How do we respond under pressure? | Crisis plan, comms protocols | Activation layer |
| Operational resilience framework | Can we stay within impact tolerances? | Tolerance statements, scenario tests, self-assessments | Integrating layer |
The business continuity plan remains the operational document teams execute during an incident. The resilience framework is the governance structure that proves the plan will work before the incident happens.
For teams still working through where business continuity ends and operational resilience begins, business resilience vs business continuity: the difference explained provides useful grounding.
Building the framework: a sequenced approach
Organizations new to this work tend to over-engineer the first iteration. A more effective approach moves in four stages:
Stage 1: Service scoping. Start with a short list. Identify five to ten services with the highest potential for customer or market harm. Resist the urge to catalogue everything; that comes later.
Stage 2: Tolerance-setting workshops. Bring together business owners, risk, legal, and technology leads for structured workshops. Anchor tolerances in real harm, not gut feel. What does two hours of payment downtime actually cost a customer? What does eight hours cost?
Stage 3: Dependency mapping. For each scoped service, trace the full chain of people, technology, facilities, and third parties. The Snowflake customer breaches of mid-2024, in which threat actors accessed data from Ticketmaster, Santander, and others through compromised credentials, showed that dependency chains extend to SaaS providers and identity infrastructure that many mapping exercises miss.
Stage 4: Testing and iteration. Run a scenario against one or two services in the first cycle. Treat the results as learning, not audit findings. Refine the map, adjust tolerances where evidence justifies it, and broaden scope in subsequent cycles.
Common gaps that auditors find
Based on regulatory findings published by the FCA and PRA, and incident post-mortems from firms that experienced operational failures, several gaps recur:
- Impact tolerances set by IT teams without business sign-off, making them RTO proxies rather than customer harm thresholds.
- Dependency maps that stop at the first-tier supplier without capturing sub-processors.
- Scenario tests that exclude scenarios the firm has actually experienced, because internal politics make live incidents uncomfortable to revisit.
- Board self-assessments signed off without evidence that the board reviewed underlying test results.
The BCM maturity benchmark 2026 offers external reference points for where most organizations sit across these dimensions.
Technology's role in maintaining the framework
A resilience framework that lives in spreadsheets and shared drives erodes quickly. Service maps become stale. Tolerance statements lose their links to supporting evidence. Testing logs sit in email threads.
Platforms that integrate BIA data, dependency mapping, and testing workflows keep the framework current between annual reviews. For teams evaluating options, BCM software: a buyer's guide covers what to look for, and how to choose business continuity software provides a practical checklist.
Automation is increasingly available for specific tasks. How to automate business continuity testing with AI covers where automation adds genuine value versus where human judgment remains essential.
What good looks like
A mature operational resilience framework produces a small number of outcomes that are easy to verify: the organization knows which services are critical and why, each service has a board-approved impact tolerance, the dependency map is current and includes third parties, testing has been run against each tolerance in the past twelve months, and the board has reviewed and approved the annual self-assessment.
That sounds simple. Most organizations are not there yet.

