Cybersecurity: A Pillar of Operational Resilience cover
Download the PDF

PDF · opens in a new tab

// white paper

Cybersecurity: A Pillar of Operational Resilience

The Role of Governance in Implementing and Sustaining Robust OT Security Programs

Chuck Toomey, P.E. ·

OT resilience is the ability to keep essential plant functions during and after a cyber incident. Governance is who is allowed to make that call, against which risk, with which evidence. Tools do not make it. A RACI, a BIA that treats IT and OT as different failure modes, and a tested escalation path do.

Colonial Pipeline in May 2021 is the US case. DarkSide hit IT. OT was shut as a precaution because the governance framework did not say how to run the pipeline while IT was down. Fuel shortages followed. Decide the IT-hit / OT-up call before the incident. Segmentation without that call still stops the line.

What the working stack says

NIST CSF 2.0 Govern. Oversight, policy, and accountability as a function, not a preface. Cross-functional review of controls belongs in the operating rhythm.

ISO 27001. ISMS: policy, roles, internal audit, management review. Useful for the corporate loop. It will not, by itself, write the OT shutdown protocol.

IEC 62443. Roles and security policy across the IACS lifecycle, from design through operation. 62443-3-2 is the risk-assessment and zone-design clause plants can run.

NIST SP 800-82r3. ICS-specific incident response, playbooks, and recovery tied to operational objectives. This is the document to put in the IR team’s hands.

CMMI. A maturity lens on process discipline. Use it to score whether governance is a binder or a habit.

The PDF compares those five on governance, risk, policy, metrics, leadership, and incident response. Steal the table. Do not flatten them into one “framework program.”

Two incidents, same gap

Colonial Pipeline (2021). No predefined IT-hit / OT-up protocol. Weak BIA on IT-OT interaction. Leadership decisions lagged because the decision rights were not on paper. Result: a full operational shutdown that the payload did not require.

Ukraine grid (2015 and 2016). Spear-phish into IT, then SCADA. About 225,000 customers in 2015. Attackers present for months. 2016 used Industroyer against grid stability. Failures: coordinated IT/OT response, detection, audits that would have caught access-control gaps, and communication that lets operators switch to manual without a committee.

What to install

  1. A governance structure that names IT, OT, safety, and the executive who can keep the plant up. OT governance operationalization is that work.

  2. Continuous risk ranking so compensating controls and exceptions have an owner. IEC 62443-3-2 is the method.

  3. Incident-response governance you have exercised. OT incident response planning and simulation is the drill, not the binder.

  4. Communication paths that survive the first hour. Write them. Then fail them in a tabletop.

If governance still dies at the department boundary, read OT governance has to cross security, engineering, operations, and safety.

Chuck Toomey, P.E., wrote the paper (September 2024). The PDF carries the comparison table and the incident write-ups. This page is the citable argument.

// white paper FAQ

Questions this paper answers

Why does OT resilience depend on governance?
The shutdown call is a decision right. Tools do not make it. When IT is hit and OT is still up, someone has to know whether the plant keeps running. If roles, escalation, and a pre-agreed BIA are missing, leadership shuts the line out of caution. That is what happened at Colonial Pipeline in 2021.
Which frameworks assign OT governance?
NIST CSF 2.0 added the Govern function. IEC 62443 assigns roles across the IACS lifecycle. NIST SP 800-82r3 ties ICS incident response to operational objectives. ISO 27001 gives the ISMS policy-and-review loop. CMMI is a maturity lens on those processes. Use the clause, not a generic stack.
What did the 2015-2016 Ukraine grid attacks show?
Attackers lived in the networks for months, then used SCADA to open breakers. About 225,000 customers lost power in 2015. The failures were detection, IT/OT response coordination, and tested decision rights. Industroyer in 2016 aimed at grid stability, not only outage hours.
Where do you start if IT, OT, and the board do not share a RACI?
Name who can keep the plant up during an IT incident, write the protocol, and exercise it. Hillstrong's OT governance operationalization and incident-response planning work is that job. Do not wait for the next audit to discover the gap.

Turn the research into a running program

Book a 30-minute demo. We will show how Resilion turns your existing assessments into a program you can run and prove.