An Approach To Disaster Recovery In OT
Saltanat Mashirova
S4x24 - ICS Security Conference · Day 3 · Stage 3
Overview
In the realm of Operational Technology (OT), where the convergence of physical and digital systems can have profound safety and operational consequences, the ability to recover from a cyber incident is paramount. Saltanat Mashirova's talk, "An Approach To Disaster Recovery In OT," at S4, delves into a structured methodology for enhancing disaster recovery (DR) capabilities within industrial environments. Moving beyond the conventional focus on backups and restoration, Mashirova advocates for a comprehensive approach that integrates lessons from process safety to define, classify, and respond to cyber-induced disasters with greater efficacy.

Key moments
- 0:00 Introduction to disaster recovery concepts and elements
- 2:15 Speaker's focus: loss scenarios & optimal recovery sequence
- 2:50 Using process safety to define disaster criticality
- 4:07 Cyber security shutdown levels based on safety analogy
- 7:35 Key takeaways: Disaster recovery is more than backups
- 8:50 Clarifying OT disaster recovery vs. business continuity
An Approach To Disaster Recovery In OT
Speakers: Saltanat Mashirova
Conference: S4
YouTube: https://www.youtube.com/watch?v=zjwUwGa3rLw
Overview
In the realm of Operational Technology (OT), where the convergence of physical and digital systems can have profound safety and operational consequences, the ability to recover from a cyber incident is paramount. Saltanat Mashirova's talk, "An Approach To Disaster Recovery In OT," at S4, delves into a structured methodology for enhancing disaster recovery (DR) capabilities within industrial environments. Moving beyond the conventional focus on backups and restoration, Mashirova advocates for a comprehensive approach that integrates lessons from process safety to define, classify, and respond to cyber-induced disasters with greater efficacy.
The presentation underscores the critical distinction between routine cyber incidents and true disasters that necessitate a full-scale recovery effort. By drawing a parallel with established Emergency Shutdown (ESD) systems in process safety, Mashirova proposes a tiered framework for categorizing cyber incidents based on their criticality and potential impact on operations. This innovative perspective aims to optimize the reconstitution time of production processes by developing optimal recovery sequences tailored to specific loss scenarios, thereby fortifying industrial resilience against increasingly sophisticated cyber threats.
This article provides a detailed exploration of Mashirova's proposed framework, examining the underlying principles, technical specifics, and practical implications for OT security practitioners. It highlights the importance of proactive planning, cross-functional collaboration, and rigorous testing in building robust disaster recovery capabilities that ensure industrial systems can "get back up" swiftly and safely after being "knocked down" by a cyber attack.
Background
▶ Watch: Introduction to disaster recovery concepts and elements (0:00)
The concept of disaster recovery in OT environments is often narrowly perceived as the mere act of backing up data and restoring systems. However, as Mashirova highlights, a true disaster recovery process is far more intricate, invoked by an emergency response team following an interruption of normal production. This process spans from the initial activation and notification phase, where disasters are announced and assessed, through the recovery phase, focused on restoring technical functions, and finally to the reconstitution phase, which involves restarting production.
To frame the discussion, Mashirova introduces the widely recognized bow tie diagram for cyber attack scenarios. On the left, cyber threat actions by actors like insider threats, activists, or criminals lead to cyber incidents (e.g., unauthorized access to engineering workstations to manipulate configurations). Proactive barriers (prevent and detect) aim to mitigate these threats. However, when these barriers fail, the focus shifts to the right side of the bow tie: consequences such as loss of view, loss of control, or shutting down process controllers. Here, safeguards come into play, representing an organization's resilience—its ability to respond and recover. This is where disaster recovery becomes critical.
Traditional disaster recovery planning relies on key metrics like Recovery Time Objective (RTO), the maximum tolerable downtime; Recovery Point Objective (RPO), the maximum tolerable data loss; and Maximum Tolerable Downtime (MTD). While these are essential, Mashirova argues that the conventional approach often overlooks the nuanced aspects of loss scenarios and the development of an optimal recovery sequence to minimize reconstitution time. The core challenge in OT is that a cyber incident can have direct physical consequences, impacting safety, environment, and production in ways that differ significantly from typical IT disruptions.
To address this complexity, Mashirova proposes leveraging the established methodologies of process safety, particularly the tiered approach of Emergency Shutdown (ESD) systems. Process safety, by its nature, is designed to protect humans from machines by initiating measures to reduce hazards like fire, toxic exposure, or explosions, often by shutting down equipment or entire plants. This mature framework, with its predefined states and criticality levels, offers a robust model for classifying and responding to cyber incidents that could trigger similar operational disruptions. The speaker posits that by adapting this well-understood framework, OT environments can develop more sophisticated and effective disaster recovery plans that go beyond simple data restoration.
Key Findings
▶ Watch: Using process safety to define disaster criticality (2:50)
The central contribution of Mashirova's talk is the innovative application of process safety's tiered Emergency Shutdown (ESD) system methodology to classify and manage cyber-induced loss scenarios in OT. This analogy provides a structured, criticality-based framework for defining what constitutes a "disaster" in a cyber context and subsequently how to orchestrate an effective recovery.
In process safety, ESD systems are categorized by the severity of the deviation from normal operating limits:
- ESD3: Deviation outside operating limits of a package or equipment (process upsets, well-contained incidents, not immediately threatening personnel safety).
- ESD2: Deviation outside operating limits of a process unit (similar to ESD3, but broader scope).
- ESD1: Emergency situation requiring the plant to be brought to a predefined safe state.
- ESD0: Decision to abandon the facility, representing the highest criticality.
Mashirova directly maps these process safety levels to cyber security shutdown levels, recognizing that while "process safety protects human from machines, cyber security protects machines from the human," the underlying principles of criticality and response can be adapted. This mapping defines different levels of cyber incidents that invoke disaster recovery:
- Shutdown 3 (SD3): An anomaly incident that is local and does not cause process downtime. This corresponds to a contained incident, similar to ESD3 where a single piece of equipment might be affected.
- Shutdown 2 (SD2): A noteworthy cyber incident that is still local but results in the loss or impairment of support functions and some loss of ability to perform. Crucially, the process downtime is less than the Recovery Time Objective (RTO). This signifies a more impactful local event, akin to an ESD2 affecting a process unit.
- Shutdown 1 (SD1): This level, though not explicitly detailed in the brief transcript, would logically correspond to an emergency cyber situation requiring the entire OT plant or a significant portion to be brought to a predefined safe state, mirroring ESD1. This implies a broader, more severe cyber attack necessitating a significant operational halt and coordinated recovery.
- Shutdown 0 (SD0): Analogous to ESD0, this would represent a catastrophic cyber incident requiring the abandonment of the facility or a complete, prolonged shutdown and extensive rebuilding of the OT environment.
The key finding here is that a "larger disaster recovery" is "so much more than restoration from backups." Even with a robust 3-2-1 backup strategy and regular restoration practice, a severe incident demands more. The framework emphasizes the necessity of developing:
- Recovery Sequence: A predefined, step-by-step plan for restoring operations.
- Automation Functions: Leveraging automation to expedite recovery steps where possible.
- Recovery Priorities and Dependencies: Clearly identifying which systems and functions are critical and their interdependencies, ensuring a logical and efficient restart.
- Reconstitution of the Production Process: The final, often overlooked, step of safely and efficiently bringing the plant back to full operational capacity, which is distinct from merely restoring data or systems.
By adopting this tiered, process-centric view, organizations can move beyond generic recovery plans to create targeted, optimized strategies that account for the specific nature and severity of cyber incidents, ultimately reducing the reconstitution time and improving overall resilience.
Technical Deep Dive
▶ Watch: Cyber security shutdown levels based on safety analogy (4:07)
The technical core of Mashirova's approach lies in the structured classification of cyber incidents using the ESD (Emergency Shutdown) system analogy and the subsequent development of optimal recovery sequences. This methodology aims to bring the rigor of process safety engineering to the often less mature field of OT cyber disaster recovery.
Let's dissect the proposed cyber shutdown levels and their technical implications:
- Shutdown 3 (SD3) - Anomaly Incident / Local Incident:
- Description: This level represents a localized cyber incident that causes an anomaly but does not lead to immediate process downtime. Examples might include a single engineering workstation compromised, a non-critical HMI experiencing a temporary glitch due to malware, or a minor network segment disruption that doesn't propagate.
- Technical Impact: While there might be some loss of view or minor control issues on a specific asset or package, the overall process continues. The incident is contained.
- Recovery Focus: Localized remediation. This could involve isolating the affected asset, cleaning malware, restoring configuration from a local backup, or re-imaging a single workstation. The recovery sequence here is relatively simple and contained within a specific operational domain, often handled by local OT support teams.
- Analogy to ESD3: A deviation outside operating limits of a single piece of equipment or a package, manageable without broader plant shutdown.
- Shutdown 2 (SD2) - Noteworthy Cyber Incident / Impaired Support Function:
- Description: This signifies a more significant local incident where support functions are impaired, and there's a demonstrable loss of ability to perform certain operations. Crucially, while there is process downtime, it is less than the Recovery Time Objective (RTO). This suggests that the organization has a predefined acceptable downtime, and this incident falls within that window.
- Technical Impact: Could involve a small cluster of HMIs, a control server, or a specific control network segment impacting a process unit. Loss of historical data, impaired alarm management, or degraded operator control for a specific area might occur. While production might be affected, a full plant shutdown is not yet necessary.
- Recovery Focus: This requires a more coordinated recovery effort, potentially involving multiple teams (OT, IT, vendor support). The recovery sequence would involve restoring affected servers, network devices, and potentially re-establishing communication paths. Prioritization becomes important to restore essential support functions first. The goal is to bring the affected process unit back online within the RTO.
- Analogy to ESD2: A deviation outside operating limits of a process unit, requiring intervention but not a full plant emergency.
- Shutdown 1 (SD1) - Emergency Cyber Situation / Predefined Safe State:
- Description: Although not explicitly detailed in the provided transcript, consistent with the ESD analogy, SD1 would represent a severe cyber incident demanding that the entire plant or a major process area be brought to a predefined safe state. This implies a direct threat to safety, environment, or widespread operational integrity.
- Technical Impact: This level could be triggered by widespread malware (e.g., ransomware affecting multiple control systems), a loss of critical safety functions, or an attacker gaining pervasive control over industrial processes. Loss of view and control across significant portions of the plant would be expected.
- Recovery Focus: A full-scale emergency response. The recovery sequence would be highly complex, involving a systematic shutdown, isolation of affected systems, forensic analysis, complete restoration of critical control systems (DCS, PLC, SIS), network infrastructure, and then a carefully orchestrated, step-by-step restart of the entire production process. This would involve rigorous testing at each stage to ensure safety and functionality. This is where recovery priorities and dependencies are paramount, as an incorrect sequence could lead to further hazards or delays.
- Analogy to ESD1: An emergency situation requiring the plant to be brought to a predefined safe state.
- Shutdown 0 (SD0) - Decision to Abandon Facility:
- Description: The most catastrophic level, where the cyber incident is so severe, widespread, and irrecoverable in the short-to-medium term that it necessitates a decision to abandon the facility (temporarily or permanently).
- Technical Impact: This would imply complete compromise and destruction of critical OT infrastructure, inability to regain control, or an incident that renders the facility unsafe or economically unviable to operate.
- Recovery Focus: Long-term strategic planning, potentially involving rebuilding significant portions of the OT environment from scratch, or even decommissioning the facility.
- Analogy to ESD0: Decision to abandon the facility due to insurmountable hazards.
The development of an optimal recovery sequence for each of these identified loss scenarios is the cornerstone of this approach. This isn't just a list of steps; it's a meticulously planned choreography of actions, considering:
- System Interdependencies: Understanding which systems must be online before others (e.g., network infrastructure before control servers, control servers before PLCs, safety systems before process controllers).
- Process Requirements: The specific operational states and sequences required to safely restart a complex industrial process. This is where the deep knowledge of process people (engineers, operators) is indispensable.
- Automation Opportunities: Identifying tasks that can be automated to reduce human error and accelerate recovery (e.g., automated configuration deployments, system re-imaging).
- Verification Steps: Built-in checks at each stage to ensure successful restoration and safe operation before proceeding to the next step.
Finally, the emphasis on reconstitution of the production process highlights that "getting up" means not just functional systems but a fully operational and safe plant. This phase involves bringing back production, validating product quality, and ensuring stable operation, which requires a holistic view beyond just IT-centric system recovery. The technical deep dive reveals that effective OT disaster recovery is a complex engineering challenge, not merely an IT administrative task.
Demo / Proof of Concept
▶ Watch: Key takeaways: Disaster recovery is more than backups (7:35)
The talk by Saltanat Mashirova is primarily conceptual and methodological, focusing on a framework for classifying cyber-induced disasters and structuring recovery efforts in OT environments. As such, it does not include a live demonstration or a proof of concept of a specific tool or technology. Instead, the presentation outlines a strategic approach, drawing analogies from process safety to build a more robust and effective disaster recovery plan. The value lies in the proposed methodology and the intellectual framework for addressing complex OT recovery challenges, rather than a hands-on technical demonstration.
Defensive Implications
▶ Watch: Clarifying OT disaster recovery vs. business continuity (8:50)
Saltanat Mashirova's framework offers several critical defensive implications for organizations operating in the OT space, urging a shift from reactive, IT-centric recovery to proactive, process-informed resilience.
- Define and Tier Cyber Disasters: The most immediate implication is the need for organizations to clearly define what constitutes a "cyber disaster" within their specific OT context. By adopting a tiered approach similar to the ESD levels (SD3, SD2, SD1, SD0), defenders can develop targeted response plans proportionate to the incident's criticality. This moves beyond a binary "incident/no incident" view to a nuanced understanding of impact, allowing for appropriate resource allocation and escalation.
- Develop Optimal Recovery Sequences: Beyond simply having backups, OT environments must meticulously develop optimal recovery sequences for each identified loss scenario. This involves:
- Mapping Dependencies: Identifying critical systems, their interdependencies, and the logical order in which they must be restored. This includes control systems (DCS, SCADA), safety instrumented systems (SIS), network infrastructure, and supporting IT systems.
- Prioritization: Clearly defining which functions and assets are absolutely essential for safe operation and must be recovered first. This is where essential functions from the process perspective become the guiding principle, not just IT system uptime.
- Automation: Exploring opportunities to automate recovery steps to reduce manual error, accelerate response times, and ensure consistency.
- Involve Process People: A recurring theme is that disaster recovery in OT "cannot be done by IT team" alone. It requires deep collaboration with process people—plant operators, process engineers, and maintenance staff. These individuals possess invaluable knowledge about plant functions, operational priorities, and the safe startup sequences, which are vital for identifying essential functions and building realistic recovery plans. DR teams should include individuals responsible for network levels 3.5 (firewall) and down to the process control network.
- Practice Beyond Backups: While a 3-2-1 backup strategy is foundational, it's insufficient. Defenders must regularly practice the full restoration and reconstitution process, not just data recovery. This includes testing the entire recovery sequence, validating system functionality post-restoration, and ensuring the safe restart of production.
- Create Worst-Case Scenarios: Organizations must proactively "create your worst case cyber attack scenarios" to prepare effectively. This involves threat modeling and tabletop exercises that simulate severe incidents to expose weaknesses in recovery plans and build muscle memory for response teams.
- Strategic Testing of Recovery Plans: Testing disaster recovery plans in a live OT environment is challenging, if not impossible, for critical infrastructure. Mashirova suggests several strategies:
- Greenfield Projects: During the commissioning phase of new facilities (before "ready for operation"), conducting a black start test is ideal.
- Virtualization Test Beds: Utilizing virtualization to create realistic test environments that mimic the OT network and control systems allows for safe and repeatable testing of recovery sequences without impacting live operations.
- Turnaround Times: Leveraging planned maintenance turnaround times or scheduled shutdowns to conduct limited, controlled tests of recovery procedures.
- Brownfield Challenges: Acknowledge that comprehensive black start testing in existing (brownfield) environments is "almost near to impossible," underscoring the importance of greenfield testing and virtualization.
By embracing these defensive implications, OT organizations can evolve their disaster recovery capabilities from a reactive measure to a proactive, integrated component of their overall cyber resilience strategy, ensuring the safety and continuity of critical industrial operations.
Key Takeaways
- DR in OT is More Than Backups: Effective disaster recovery in Operational Technology extends far beyond simple backup and restoration; it requires a holistic approach encompassing incident classification, recovery sequencing, and production reconstitution.
- Leverage Process Safety Methodologies: Applying the structured, tiered approach of Emergency Shutdown (ESD) systems from process safety to classify cyber incidents (SD3, SD2, SD1, SD0) provides a robust framework for defining disaster criticality and guiding response.
- Optimal Recovery Sequences are Critical: For each identified cyber loss scenario, organizations must develop detailed, prioritized, and dependency-aware optimal recovery sequences to minimize reconstitution time and ensure safe operational restart.
- Cross-Functional Collaboration is Essential: Disaster recovery planning in OT is not solely an IT responsibility; it mandates deep collaboration with process people (engineers, operators) who understand plant functions, operational priorities, and safe startup procedures.
- Proactive Scenario Planning and Testing: Organizations must create and simulate worst-case cyber attack scenarios. Testing should leverage greenfield projects, virtualization test beds, or planned turnaround times, acknowledging the difficulty of testing in live brownfield environments.
- Focus on Reconstitution, Not Just Restoration: The ultimate goal of disaster recovery is the safe and efficient reconstitution of the production process, meaning bringing the plant back to full operational capacity, which goes beyond merely restoring systems or data.
About the Speaker(s)
Saltanat Mashirova is a speaker at the S4 conference, where she presented her insights on "An Approach To Disaster Recovery In OT." Based on her talk, she is an expert in industrial cybersecurity and operational technology, with a focus on developing resilient strategies for critical infrastructure. Her expertise lies in bridging the gap between traditional IT disaster recovery practices and the unique requirements and challenges of OT environments, particularly by integrating concepts from process safety.
Reviews
Dr. Zero (Offensive Security Researcher) — MUST SEE
Mashirova's talk delivers a critical paradigm shift for disaster recovery in Operational Technology. By ingeniously leveraging the mature, tiered framework of process safety's Emergency Shutdown (ESD) systems, she provides a robust, criticality-based methodology for classifying cyber incidents and orchestrating optimal recovery sequences. This isn't another rehash of 3-2-1 backups; it's a profound, actionable approach that demands deep collaboration between cyber and process teams, ultimately fortifying industrial resilience against catastrophic cyber events. This is the kind of practical, first-principles thinking this industry desperately needs.
Heather Calloway (CISO) — STRONG ACCEPT
Saltanat Mashirova's talk offers a critical reframing of disaster recovery in Operational Technology, moving beyond mere backups to a comprehensive, tiered approach. By drawing a powerful analogy to process safety's Emergency Shutdown (ESD) systems, the presentation provides a robust framework for classifying cyber-induced disasters and developing optimal recovery sequences. This strategic shift demands cross-functional collaboration and a keen focus on business reconstitution, making it highly relevant for security leaders and executives grappling with real-world OT resilience.