A Useful OT RACI Starts With Decision Rights, Not Departments
Hillstrong Group Security ·
A RACI that looks clean in PowerPoint usually fails in the first real conflict.
That is not because RACI is useless. It is because most companies use it to map functions instead of decisions.
Security gets a box. Engineering gets a box. Operations gets a box. Somebody adds safety, procurement, and risk because the matrix looks mature with more names on it. Then a real OT issue shows up and the model falls apart.
A fragile controller cannot be patched before peak production.
A vendor needs remote access during an outage.
A segmentation change will improve the standard but could disrupt restart timing.
A ransomware event forces a choice between immediate isolation and continued production under uncertainty.
Now the question is no longer who is Responsible, Accountable, Consulted, or Informed in theory. The question is who decides.
That is where most OT RACIs fail.
The underlying research base for this series makes the point in a blunt, usable way. A useful OT RACI should be built around decision points, not generic departments.[1] The strongest decision areas are straightforward:
- who owns enterprise OT standards
- who owns the site-level implementation plan
- who approves OT exceptions
- who has authority during a production-impacting incident
- who sponsors and governs third-party OT access[1]
That list matters because it moves the conversation away from abstract organizational design and into operating reality. A manufacturing CISO does not need another matrix that describes the organization at rest. The enterprise needs a decision-rights model that works when there is money on the line, production at risk, and no clean answer.
The problem with most OT RACIs
Most OT RACIs fail for one simple reason. They assign people to nouns instead of verbs.
They map departments to categories like patching, segmentation, incident response, or vendor management. That sounds fine until a local reality cuts across all of them.
Take a patch deferment request on a line controller that drives a critical process cell.
Security says patching sits under vulnerability management.
Engineering says the controller is too fragile for a normal test cycle.
Operations says the next downtime window is six weeks away.
Safety says any change requires validation because a failed restart could affect process behavior.
Procurement says the OEM support agreement is fixed and onsite validation will cost extra.
The RACI still looks complete.
The decision is still stuck.
That is because the matrix described ownership of a function. It never described ownership of the decision when the function collides with production and safety.
A useful OT RACI starts somewhere else. It asks, “What are the few decisions that keep breaking governance when the easy answer disappears?”
The source base supports five of them.[1][2][3][4][5]
Get those five right and a lot of the organizational noise gets smaller.
Get them wrong and the rest of the chart is decoration.
Decision one: who owns enterprise OT standards
This is the easiest one to dodge badly.
Many organizations say standards are shared. That usually means the standard will get written by whoever has the strongest process discipline and then fought over by everybody else.
The evidence base points to a stronger answer. Enterprise security leadership should own enterprise OT standards, with OT-specific input and formal consultation from operations and engineering.[1][2][5]
That answer makes sense for practical reasons.
Standards need a home that can hold consistency across sites.
They need to connect to enterprise reporting, governance, third-party access expectations, and board-visible risk themes.
They need somebody who can see whether one plant’s workaround is becoming another plant’s default design.
That is enterprise work.
The plant should shape how the standard gets implemented. The plant should not own whether the standard exists.
This is where some OT debates get childish. Security says, “We own cyber.” Engineering says, “You do not run the plant.” Operations says, “You both ignore production windows.”
All three statements can be true. They still do not answer where the standard belongs.
The standard belongs with the enterprise function that can own consistency, reporting, and escalation. The execution path belongs somewhere else.
That split is the beginning of a real RACI.
Decision two: who owns the site-level implementation plan
This one usually breaks right after the first one.
The enterprise issues a standard. Everyone agrees in principle. Then the site says the sequence will not work.
The first shutdown window is too short.
The second is tied to another maintenance activity.
The OEM has to be onsite for restart.
A process validation has to happen before the change goes live.
A line upgrade already changed the timetable.
Who owns that plan.
The evidence points to the site or system owner, working with local OT engineering and operations.[1][2][5]
That is the only answer that respects how plants actually run.
A corporate team can own the standard for segmentation, logging, remote access, or recovery testing. It should not pretend to own the site-level sequence for how that control lands on a live production system.
This distinction matters because many OT programs confuse ownership of the standard with ownership of the timetable.
That creates two bad outcomes.
The first is corporate fantasy planning. The standard looks good, the dates look clean, and the site cannot execute any of it without disrupting production.
The second is local veto by inertia. The plant says the timing is bad, nobody owns the sequence well enough to challenge or refine it, and the issue gets pushed into the next quarter.
A better model is blunt.
Corporate owns the standard.
The site owns the implementation plan.
The site still has to explain the plan in operational terms, not in excuses. That means stating the maintenance window, the process constraint, the sequencing risk, the resource need, and the earliest credible path.
Now the enterprise has something to govern.
Without that local ownership, the enterprise is asking for compliance without understanding the operating path.
Decision three: who approves OT exceptions
This is where companies reveal whether the RACI is real.
Most OT environments carry exceptions. Unpatchable assets. Temporary connectivity allowances. Deferred segmentation changes. Unsupported systems with compensating controls. Vendor access that violates the default rule but remains necessary to keep the process running.
The issue is not whether those exceptions exist. They do.
The issue is who can approve them, at what level, based on what evidence, and for how long.
NCSC guidance is direct on this point. The decision not to update is a business risk decision that should be recorded in the wider risk-management framework.[3] CISA and related OT guidance also point toward documented exceptions and compensating controls when ideal practices are not feasible.[2][4]
That means exception approval cannot live inside convenience.
It needs a named risk owner at the right level of materiality.[1][3][4]
That owner may not be the same person every time. A minor local exception with strong compensating controls may stay low. A repeated exception affecting a critical process family across multiple plants may need much higher review.
What matters is that the organization knows who decides.
A weak OT RACI says, “Operations and security will align on exceptions.”
That is not a decision path. That is a hope.
A better model says the system owner documents the issue, the local operating constraint, the proposed compensating controls, and the business impact. The appropriate risk owner decides whether to accept, reject, time-bound, escalate, or fund the root fix.
That is governance.
This is also where a lot of organizations lose credibility with plant leaders. They ask sites to request exceptions through a process that ignores plant reality, then act surprised when sites work around it informally. If you want a real exception process, the RACI has to support it.
Decision four: who has authority during a production-impacting OT incident
This is the hardest decision in the whole model.
The source base is clear that incident authority needs to be defined in advance, even though it does not prescribe one universal answer.[1][5] That is enough to make the core point. If the company cannot answer who can isolate, shut down, continue degraded operations, or sequence recovery during a production-impacting incident, the RACI is decorative.
This is where bad governance gets exposed in public.
A ransomware event touches IT and raises concern about OT spread.
Security wants to isolate now.
Operations wants to keep the line running until there is evidence of impact.
Engineering wants to preserve system stability.
Safety wants to understand what degraded mode will do to the process.
Executives want one answer.
The worst possible time to decide who has authority is in that room.
That answer has to exist before the event.
It does not need to be simplistic. In fact, it should not be. Different incident severities and different process contexts may drive different decision paths. A packaging line and a high-hazard batch process should not be governed with cartoon simplicity.
Still, the model needs to answer a few non-negotiable questions.
Who can authorize immediate isolation.
Who can authorize continued production under elevated risk.
Who can approve degraded operations.
Who owns recovery sequencing when security urgency and production urgency pull in different directions.
Who speaks for safety consequences.
SANS makes this real from the incident side. Roles and responsibilities have to be worked out across engineering, operations, safety, OT security, and IT security before the event tests them.[5]
That is why a useful RACI does not stop at incident response as a category. It assigns authority for the specific choices that create loss, recovery, or escalation under pressure.
Decision five: who sponsors and governs third-party OT access
Weak OT governance models break on vendor access faster than anything else.
That is because third-party access sits right in the seam between enterprise control and local necessity.
The plant needs the OEM.
The vendor needs speed.
The enterprise needs consistency, logging, approval logic, and a trust model it can defend across the footprint.
The evidence base supports centralized governance plus local sponsorship for third-party OT access because vendor access expands the trust boundary and often crosses organizational lines.[1][2][4]
That sentence should drive a lot more RACI design than it usually does.
The enterprise needs to own the control model. The access method, approval expectations, identity pattern, monitoring, and review path cannot be left entirely to local habit.
The site still needs to sponsor the access because the site knows when it is operationally necessary, who requested it, and what system it touches.
That split matters.
If enterprise controls third-party access with no local sponsor, the model becomes rigid and slow. Plants will route around it the minute an outage forces urgency.
If plants sponsor third-party access with no enterprise control model, the enterprise will normalize inconsistent trust boundaries across sites.
A useful RACI answers both sides.
Who owns the enterprise vendor-access standard.
Who sponsors the access locally.
Who approves exception cases.
Who reviews repeated patterns across sites.
Who owns contract and access-model drift over time.
That is how you stop the same OEM pathway from becoming five different governance models across ten plants.
What a useful OT RACI sounds like in practice
Forget the matrix for a second.
Listen to how the organization talks.
A weak RACI sounds like this:
“Security owns standards. Engineering owns systems. Operations owns uptime.”
True. Useless.
A better RACI sounds different.
“Corporate security owns the remote access standard. The site system owner owns the implementation path. If the site cannot meet the standard before the planned outage, the system owner documents the constraint and compensating controls. The plant leader and OT security lead review it. If the exception crosses the materiality threshold, the named enterprise risk owner decides whether to accept, time-bound, or escalate it. During an outage event, the incident authority model determines whether the vendor gets emergency access and who approves the recovery path.”
That is longer. It is also real.
Real governance sounds specific because real decisions are specific.
That is the standard manufacturing CISOs should bring into the room. Not an abstract RACI. A decision-rights model that survives contact with the plant.
A scene most enterprises should test against
Imagine a segmentation project on a fragile production line.
The enterprise standard says the line-support assets need tighter control boundaries.
The plant agrees in principle.
The next shutdown window is four hours shorter than expected because another maintenance task slipped.
The OEM warns that restart support will require live remote access during the cutover.
Safety says the changed sequence needs review because a failed restart could create abnormal process behavior.
Now run the RACI.
Who owns the standard.
Who owns the local implementation plan.
Who approves the temporary remote access condition.
Who decides whether the outage window is too risky for the original sequence.
Who can accept the temporary exception if the work gets deferred.
Who sees the issue if the same deferral has happened in three other sites.
If the organization cannot answer those questions quickly, the RACI is incomplete.
That is the test.
What boards and executive teams should ask
Boards do not need to review the matrix itself. They do need to know whether the organization can answer a few critical questions.
Who approves a material OT exception.
Who has authority during a production-impacting incident.
Who owns the enterprise standard when plant constraints push back.
Who owns vendor-access governance across sites.
Who decides whether a repeated local workaround has become enterprise risk.
If leadership cannot get crisp answers, the board should assume the OT RACI is not finished.
That is not a paperwork problem. It is a risk-management problem.
What this means for manufacturing CISOs
If you own the OT program, stop asking whether the RACI has enough names on it.
Ask whether it can answer five hard decisions under pressure.
Who owns the standard.
Who owns the site plan.
Who approves the exception.
Who has incident authority.
Who governs vendor access.
If your current model cannot answer those clearly, fix that before you add more governance machinery.
If your current model can answer them on paper but not in a live discussion, run the exercise again and keep going.
Plants do not need more boxes on a chart. They need clear answers before the line is down and the vendor is on the phone.
And that takes you to the last stress test.
The best place to see whether the RACI is real is the exception process, especially around vendor access and connectivity.
That is the next article.
Sources
- NIST, Guide to Operational Technology (OT) Security (SP 800-82 Rev. 3) — https://csrc.nist.gov/pubs/sp/800/82/r3/final
- CISA, Cross-Sector Cybersecurity Performance Goals (CPG 2.0) — https://www.cisa.gov/cybersecurity-performance-goals-2-0-cpg-2-0
- NCSC, Vulnerability management guidance (Principle 4: The organisation must own the risks of not updating) — https://www.ncsc.gov.uk/collection/vulnerability-management/guidance
- NCSC, Principle 1: Balance the risk and opportunities — https://www.ncsc.gov.uk/collection/operational-technology/secure-connectivity/principle-1
- SANS Institute, ICS/OT Cybersecurity Leadership: The Mission is Safety and Fundamentals First — https://www.sans.org/blog/icsot-cybersecurity-leadership-mission-safety-fundamentals-first