OT Exceptions, Vendor Access, and Connectivity Expose Whether Governance Is Real
Hillstrong Group Security ·
Do not judge an OT program by its slide deck. Judge it by what happens when the plant cannot patch, the OEM needs remote access, and the shutdown window is gone.
That is where the truth shows up.
Most OT organizations say they have governance. Plenty of them do. On paper.
They have standards. They have steering meetings. They have roles. They have architecture diagrams. They have exception forms. They have somebody who can explain the target state in a conference room.
Then a fragile controller reaches production season without a patch path. A vendor needs remote access for a recovery step nobody planned for. A line expansion adds connectivity faster than governance can catch up. A compensating control exists in the ticket and nowhere else. That is when you find out whether the organization has a governance model or a filing system.
The evidence base behind this series is unusually strong on this point. OT exceptions should not live in email, local custom, or tribal knowledge.[1][2][3] NCSC guidance says the decision not to update is a business risk decision that should be recorded in the wider risk-management framework.[2] NCSC OT connectivity guidance ties connectivity decisions to business-case logic, acceptable risk, and a named senior risk owner.[3] CISA and related OT guidance also support documented exceptions and compensating controls when ideal practices are not feasible.[1]
That is not administrative advice. It is the clearest governance stress test in the whole OT program.
A plant can claim to follow the standard.
A steering committee can claim to oversee the standard.
A RACI can claim to assign the standard.
The exception process tells you whether any of that survives contact with the real operating environment.
OT exceptions are normal
This needs to be said first because too much OT writing still talks as if exceptions are a sign of failure by themselves.
They are not.
They are a sign that OT is attached to real physical processes, real shutdown windows, real maintenance constraints, real OEM dependencies, and real safety consequences.
Legacy systems are normal.
Fragile assets are normal.
Production timing is normal.
OEM dependencies are normal.
Safety constraints are normal.
Connectivity inherited from modernization projects is normal.
A plant that cannot patch a controller during peak run may be making the right local call.
A site that needs temporary vendor access during a complex restart may be making the right local call.
A line team that delays a network change because the sequence is too risky before a seasonal production push may be making the right local call.
The existence of the constraint is not the indictment.
The governance question is what happens next.
Does the issue get documented with enough clarity for the enterprise to understand it.
Does someone own the risk.
Does the exception have a review path.
Do repeated patterns rise above the plant.
Does the enterprise learn anything from the fact that the same compromise is appearing in multiple places.
That is where mature and immature OT programs separate.
Normal does not mean informal
This is the center of the whole article.
A lot of manufacturers treat OT exceptions as operational facts instead of governed risk decisions.
The plant says the patch cannot happen.
The OEM says remote access has to stay open for support.
The maintenance lead says the outage window will not hold the change.
Security says the issue is understood.
Then the organization moves on.
That is not governance. That is a verbal accommodation.
NCSC says the decision not to update should be treated as a business risk decision and reflected in the wider risk-management framework.[2] That sentence deserves more attention than it gets. It means non-updating is not merely a technical delay. It is a decision with business ownership and traceable consequence.
The same logic applies more broadly.
If a vendor needs a remote path that violates the normal access model, that is not only an engineering convenience. It is a trust-boundary decision.
If a modernization project adds connectivity between systems that used to be isolated, that is not only a design choice. It is an enterprise risk decision.
If a plant cannot follow the segmentation timetable because of process constraints, that is not only a scheduling issue. It is a risk acceptance decision waiting for somebody to name it.
The difference between mature and immature OT governance is not whether those moments occur. It is whether the company treats them formally enough to learn from them, escalate them, and fund the root fix when the pattern repeats.
What a bad exception process looks like
Bad exception handling is easy to recognize once you stop being polite about it.
The first sign is local email chains.
The exception lives in messages between a plant engineer, a security lead, and a vendor contact. The reasoning may be sound. The governance is nonexistent. Six months later nobody can say who approved it, which compensating controls were supposed to exist, or when the issue was supposed to come back for review.
The second sign is undocumented vendor access.
The OEM still connects because operations needs them and everyone knows the relationship matters. The local team can explain why the access exists. Nobody can explain whether the enterprise recognizes the same pattern across the other sites that use the same vendor.
The third sign is no named risk owner.
Everyone is involved. Nobody is accountable. Security knows the issue. Engineering knows the constraint. Operations knows the business impact. No one owns the acceptance of the residual risk.
The fourth sign is compensating controls that exist in the ticket but not in the plant.
A document says the local firewall rule was tightened, logging was increased, or remote access is limited to specific windows. The configuration never changed, the log never gets reviewed, or the local team assumed someone else handled it.
The fifth sign is repeated exceptions with no enterprise view.
Plant 2 has the issue. Plant 5 has the same issue. Plant 9 uses a slightly different explanation for the same issue. Nobody aggregates the pattern. Nobody asks whether the common cause needs funding, architectural change, or a vendor reset.
That is what weak governance looks like in the wild.
Not dramatic failure. Drift.
The enterprise slowly converts one-time accommodations into normal operating practice and never says so out loud.
What a credible exception process looks like
A credible exception process is not glamorous. It is disciplined.
It starts with a named requestor.
Someone has to put the issue on paper and describe the real constraint. Not vague phrases. The actual line, asset, timing problem, vendor dependency, safety constraint, or restart risk.
It includes a named risk owner.
Not a broad committee. Not a shared mailbox. A person or clearly defined role with the authority to accept, reject, time-bound, escalate, or require additional controls.
It includes a business justification.
Not empty language about operational necessity. Specific consequence. Missed production. Unsafe timing. Lack of validated restart support. Capital dependency. Regulatory timing. Whatever the real reason is, name it.
It includes a clear description of the operational constraint.
This is where many programs fall apart. “Plant cannot patch right now” is not a constraint description. “The line runs continuously through the quarter, the OEM requires onsite validation for restart, and the next safe outage is scheduled for week 42” is a constraint description.
It includes defined compensating controls.
Real controls. Not aspirations. If logging increases, say who reviews it. If remote access is limited, say how. If network exposure changes, say where. If a manual check substitutes for a technical control, say who performs it and when.
It includes a review date.
Not because calendars make governance mature, but because open-ended exceptions turn into permanent architecture by neglect.
It includes an escalation threshold.
The organization should know when the exception stays local and when it rises. Cross-site repetition. High-hazard process impact. Third-party concentration. Long exception age. Material business consequence. Pick the triggers and state them.
Most important, it includes enterprise visibility when patterns repeat.
That is the part too many exception processes miss. The company does not just need to know the issue exists. It needs to know when the same issue stops being local.
Vendor access deserves its own section because it breaks bad governance first
Third-party OT access sits right on the fault line between plant reality and enterprise control.
The plant needs support fast.
The vendor wants a reliable path.
The enterprise wants consistency, logging, identity control, and a trust model it can defend across sites.
Weak governance lets local urgency win every time.
The result is familiar. One site sponsors access one way. Another site uses a different method for the same OEM. A third site has an older arrangement nobody wants to touch because it still works. Local teams know the story. The enterprise does not know the pattern.
That is not a vendor management issue alone. It is a governance issue.
The source base supports centralized governance plus local sponsorship for third-party OT access because vendor access expands the trust boundary and often crosses organizational lines.[1][3][4] That is one of the most important sentences in the whole OT governance argument.
The enterprise should own the control model.
The plant should own the local sponsorship and the business case.
The exception process should connect the two.
If any one part is missing, the model breaks.
Enterprise-only control without local sponsorship creates delay and workarounds.
Local-only sponsorship without enterprise control creates inconsistent access patterns, weak logging, and no cross-site view.
A good manufacturing CISO should press on this constantly. If the same OEM supports five plants, the company should not be learning about the access model one outage at a time.
Connectivity changes should trigger governance, not just design review
This is another place where organizations hide governance problems inside technical language.
A modernization project adds a new data path.
A line upgrade needs a remote support function.
A plant wants to connect a previously local system into a broader enterprise workflow.
Everyone talks about architecture. Fewer people talk about ownership.
NCSC OT connectivity guidance is useful because it cuts through that habit. Connectivity should tie back to business-case logic, acceptable risk, and named senior risk ownership.[3]
That means connectivity is not just a design matter.
It is a governance event.
Who asked for the connection.
What business value it serves.
What risk it introduces.
What control assumptions it depends on.
What the fallback plan is if the control model slips.
Who owns the residual risk.
That is what a CISO should want answered before the enterprise normalizes the new path.
A lot of manufacturing cyber debt starts this way. The connection looked practical, the project was urgent, the plant needed the capability, and nobody forced the governance discussion while the design was still flexible.
Two years later the organization treats the path as permanent and struggles to explain who approved the original risk logic.
That is exactly the kind of problem a mature exception and connectivity process is supposed to stop.
How a CISO should use exception data
Exception handling is not only a control mechanism. It is an intelligence feed about how the OT program is actually working.
Use it that way.
Repeated exceptions identify repeated control failures.
If the same exception reason keeps appearing, the problem is not local discipline. The problem is probably architectural, budgetary, or structural.
Exception data exposes underfunded modernization work.
If the organization keeps deferring controls because the underlying assets are too fragile, leadership should stop calling it a plant problem and start treating it as a capital and lifecycle problem.
Exception data surfaces vendor concentration risk.
If the same vendor access workaround appears across a cluster of critical sites, that is no longer a site support issue. It is a trust-boundary pattern the enterprise should be governing directly.
Exception data shows where plant constraints are systemic instead of local.
This is one of the biggest missed opportunities in OT governance. The plant gets blamed for resisting the standard. In reality the enterprise may have created the same impossible timing condition in multiple places.
A good CISO does not use exception data to punish plants for telling the truth. The CISO uses it to separate noise from repeated exposure.
That is where exception handling becomes strategic.
It stops being only an approval process and starts becoming a map of where the OT program is bending under real operating conditions.
A scene most manufacturers will recognize
A plant cannot patch a controller before a major production run.
The issue sounds local.
Then the review shows three other sites use the same controller family. Two have the same OEM. One has already deferred the patch twice. The compensating controls vary. Logging is consistent in one site, weak in another, and assumed in a third.
Now the question changes.
It is no longer, “Can this plant patch on time?”
It becomes, “Why does the enterprise keep finding the same unsupported path in the same device family with the same support model and no shared remediation strategy?”
That is the value of governed exception handling.
It converts a local story into an enterprise fact pattern.
That is what boards and executive teams need to see.
Not the technical detail alone. The pattern.
What boards and executive teams should ask
Boards do not need to approve routine OT exceptions. They do need visibility into the patterns that tell them whether the enterprise is carrying unmanaged OT risk.
How many material OT exceptions remain open longer than the organization expected.
Which exception categories keep repeating.
Which vendors, architectures, or asset classes show up too often.
Which issues remain unresolved because the root fix needs enterprise funding or enterprise authority.
Which connectivity changes were accepted under pressure and still have not been revisited.
If exception patterns never reach enterprise governance, the board is seeing OT risk after the decisions that create it have already been made.
That is too late.
What this means for manufacturing CISOs
If you lead OT cyber governance, stop treating exceptions as paperwork.
Treat them as evidence.
They tell you where the standard does not fit the operating reality.
They tell you where the enterprise keeps making the same compromise.
They tell you where vendors are shaping your trust boundary more than your own controls are.
They tell you where a local plant issue has already become enterprise risk.
That should change how you lead.
Push for named risk owners.
Push for review discipline.
Push for real compensating controls, not assumed ones.
Push for cross-site visibility.
Push for connectivity governance early, not after the design has already hardened.
Most of all, push for honesty. The exception process should help plants explain constraints clearly, not force them to hide reality until an incident exposes it.
That is the right place to end this series because it brings the whole argument together.
Senior visibility matters because repeated local exceptions become enterprise risk.
Named leadership and system ownership matter because someone has to speak for the standard and someone has to speak for the system.
Cross-functional governance matters because security, engineering, operations, safety, and risk all shape the decision.
Decision-rights discipline matters because a real RACI has to answer who approves the exception, who controls vendor access, and who owns the risk when the standard breaks.
The exception process is where all of those claims either survive or collapse.
The fastest way to tell whether OT governance is real is simple: look at the exception process, the vendor access model, and the last connectivity decision no one wanted to own.
Sources
CISA, Cross-Sector Cybersecurity Performance Goals (CPGs) — https://www.cisa.gov/cross-sector-cybersecurity-performance-goals-cpgs
NCSC, The organisation must own the risks of not updating — https://www.ncsc.gov.uk/collection/vulnerability-management/guidance/organisation-own-risk-updating
NCSC, Principle 1: Balance the risk and opportunities — https://www.ncsc.gov.uk/collection/operational-technology/secure-connectivity/principle-1
Cyber.gov.au, Principles of operational technology cyber security — https://www.cyber.gov.au/business-government/secure-design/operational-technology-environments/principles-of-operational-technology-cyber-security