Operational Technology in the Crosshairs: What the 2025–2026 Attacks Reveal About Industrial Cyber Risk
What the 2026 warnings from Europe and the United States reveal about exposed PLCs, cyber-physical risk, and resilient industrial architecture
An evidence-driven and formal analysis of 2025–2026 OT attacks showing how digital footholds can accumulate into operational authority and physical consequence, and deriving an IEC 62443-aligned architecture for constrained connectivity, least authority, controller integrity, detection, safe recovery, supply-chain assurance, lifecycle governance, and bounded cyber-physical compromise.
cybersecurity
energy
enterprise risk management
essay
regulation and compliance
🇬🇧
Author
Affiliation
Antonio Montano
4M4
Published
September 5, 2026
Modified
September 5, 2026
Abstract
The 2025–2026 operational-technology threat landscape shows a convergence of problems that are often analyzed separately. Internet-facing PLCs and HMIs are being actively targeted; public industrial-protocol libraries, scanning services, legitimate engineering tools, and AI-assisted scripting are reducing the effort required to interact with controllers; private carrier networks, supplier access, enterprise dependencies, and shared management infrastructure create indirect paths into systems that are not publicly exposed; and recent incidents in the United States, Poland, Denmark, Norway, the United Kingdom, and other jurisdictions demonstrate that ordinary digital privileges can become operational authority or disrupt an industrial mission without requiring novel ICS malware. The analysis therefore begins by distinguishing exposure, targeting, confirmed compromise, operational effect, and attribution, and by separating direct OT manipulation, OT impairment, and digitally induced operational disruption.
The central argument is that the primary security object in modern OT is not the individual PLC, firewall, VPN, protocol, or vulnerability, but the cyber-physical authority architecture through which identities, endpoints, trust relationships, networks, engineering systems, protocols, privileges, controller functions, safety mechanisms, and external dependencies compose. The article develops a taxonomy of sixteen failure classes across six domains (knowledge and governance, reachability and trust, identity and authority, control-path integrity, observability, and safety and resilience), and formalizes one part of that taxonomy with a typed capability-inference model. Rather than assuming a simple attack graph, the model represents conjunctive prerequisites and architecture-dependent inference rules, derives the capability closure of plausible initial compromises, distinguishes attack surface from derived operational authority, identifies high-consequence derivations and consequence concentration, and frames defensive-control selection as the problem of making unacceptable capabilities underivable even under defined control-loss scenarios.
The analysis then extends from cyber derivability to cyber-physical admissibility. Physical-process dynamics, mission viability, safe degradation, trusted control capability, external-dependency loss, safety independence, containment, and recovery are treated as engineering constraints on cybersecurity design. IEC 62443 is interpreted as a lifecycle and system-architecture framework rather than a product checklist: zones and conduits constrain reachability, Foundational Requirements constrain different parts of the authority system, and component security capability must compose with asset-owner governance, secure product development, system integration, service-provider processes, patch management, and compensating architecture. From this basis the article derives practical patterns for industrial DMZs, PAWs, brokered privileged access, ZTNA, private carrier networks, microsegmentation, unidirectional gateways, controller least authority, programming windows, engineering-workstation protection, project and firmware integrity, and secure industrial protocols including Modbus Security, CIP Security, DNP3 Secure Authentication, and OPC UA.
Prevention is treated as incomplete. The detection architecture combines context-sensitive communication baselines, passive protocol-aware observation, semantic analysis of industrial operations, controller-state verification, cross-layer identity and engineering evidence, and process-model residuals while distinguishing statistical anomaly from operational illegitimacy. Recovery is correspondingly defined as safe reconstitution rather than server restoration: containment must preserve physical safety, backups must include authoritative engineering and security state, recovery trust must remain sufficiently independent from production, controller and protection state must be verified, digital and physical process state must be reconciled, and external connectivity must be reintroduced only after high-consequence derivations remain infeasible.
The final sections move the analysis upstream and outward. Secure-by-design procurement, secure defaults, configuration governability, baseline logging, owner autonomy, secure development lifecycles, support periods, SBOMs, supplier due diligence, and cryptographic lifecycle management determine whether the required security properties remain governable over industrial lifetimes. NIS2, the EU Cyber Resilience Act, the electricity cybersecurity network code, NERC/FERC requirements, U.S. bulk-power supply-chain policy, evolving UK energy regulation, and Australia’s critical-infrastructure regime are examined as distinct legal mechanisms that increasingly assign explicit accountability for many of the same architectural properties. A reference defense-in-depth architecture is consequently derived around constrained operational authority rather than arbitrary network levels. Its objective is deliberately narrower than perfect prevention: plausible ordinary compromise should remain bounded, high-consequence authority should require increasingly specific and partially independent conditions, abnormal authority should be observable, safety and mission capability should survive defined failures, and trustworthy operation should be recoverable without reconstructing the original compromise conditions. The remaining open problems include common-mode trust, security complexity, brownfield modernization, cryptographic agility, adaptive detection, encrypted observability, controller attestation, industrial semantic authorization, AI authority boundaries, cyber-safety co-assurance, recovery assurance, and correlated fleet risk.
Keywords
operational technology, OT cybersecurity, industrial control systems, ICS, programmable logic controllers, PLC, SCADA, cyber-physical systems, critical infrastructure, operational authority, capability inference, capability closure, attack surface, defense in depth, authority containment, IEC 62443, zones and conduits, industrial DMZ, zero trust, Zero Trust Network Access, ZTNA, privileged access workstation, PAW, industrial protocols, Modbus Security, CIP Security, DNP3 Secure Authentication, OPC UA, Siemens S7, controller hardening, controller integrity, engineering workstation security, process-aware detection, OT monitoring, functional safety, safety instrumented systems, mission resilience, safe degradation, OT backup, safe reconstitution, supply-chain cybersecurity, secure by design, secure by default, SBOM, lifecycle governability, NIS2, Cyber Resilience Act, NERC CIP, critical-infrastructure regulation, AI in operational technology
An evidence-driven and formal analysis of 2025–2026 OT attacks showing how digital footholds can accumulate into operational authority and physical consequence, and deriving an IEC 62443-aligned architecture for constrained connectivity, least authority, controller integrity, detection, safe recovery, supply-chain assurance, lifecycle governance, and bounded cyber-physical compromise.
The warning is architectural, not vendor-specific
On August 19, 2026, the U.S. National Security Agency, Cybersecurity and Infrastructure Security Agency, Federal Bureau of Investigation, Department of Energy, and Environmental Protection Agency warned of an active threat to Siemens S7 programmable logic controllers. The joint advisory covers the S7-200, S7-300, S7-400, S7-1200, and S7-1500 families, including safety variants, and describes targeted reconnaissance and capability development using internet-scanning services, snap7-based tooling, S7comm, and AI-generated scripts capable of read/write interaction with controllers.1
Eight days later, the United Kingdom’s National Cyber Security Centre warned of increased targeting of operational technology across sectors and jurisdictions. The NCSC reported that some activity had already caused limited real-world disruption and urged operators to verify, rather than merely assume, that their OT environments were inaccessible from external networks. Legacy connectivity, unmanaged assets, exposed edge devices, and configuration errors can all invalidate an architecture that appears isolated on paper.2
The warning coincided with reporting of a disruptive incident involving a small-scale UK energy generator. The Register reported that a British government spokesperson had confirmed the incident and stated that the wider energy system had not been at risk. The report associated the event with suspected Iran-linked activity, while also making clear that the United Kingdom had not formally attributed it to Iran, another government, or a named hacking group.3 The publicly supportable conclusion is therefore deliberately narrow: a disruptive incident affecting a small-scale generator was government-confirmed; the victim’s identity and any state attribution remained publicly unconfirmed.
The U.S. Siemens advisory requires the same evidentiary discipline. It establishes active targeting, reconnaissance, tooling, and capability development. It describes snap7-based software able to read from and write to S7 controllers and reports observed read operations used to understand target environments and prepare for possible subsequent manipulation. It does not establish that every targeted controller was compromised, nor that the physical consequences discussed in the advisory had already occurred.4
This distinction matters because the security problem is larger than Siemens S7 and more precise than the generic proposition that PLCs are vulnerable. NIST defines operational technology as programmable systems and devices that interact with the physical environment, or that manage devices which do so. The category therefore extends beyond industrial control systems to building automation, transportation, physical-access systems, and other cyber-physical technologies.5
Within that broader category, a PLC is consequential because it occupies a privileged position in the causal chain from digital state to physical action. It samples process inputs, evaluates control logic, and produces outputs that influence pumps, valves, motors, drives, breakers, relays, and other equipment. A cyber intrusion becomes a cyber-physical problem when an attacker can interfere with that chain: by falsifying observations, changing control logic, modifying setpoints or operating modes, altering actuator commands, disabling protective functions, or making control unavailable when the process depends on it.
The important distinction is therefore between reachability and operational authority. A device may be reachable without giving the attacker meaningful control over the process. Conversely, apparently modest privileges somewhere upstream may compose with other trust relationships until they produce the ability to change a process-relevant state, or prevent the system from maintaining that state within its required safety and availability constraints. OT risk is determined not merely by whether communication is possible, but by what authority can ultimately be derived from that communication path.
The August advisory’s reference to AI-generated tooling belongs inside this model. Its significance is not that autonomous AI has suddenly made sophisticated industrial sabotage trivial. The more defensible conclusion is that part of the effort required to develop controller-specific capability is becoming cheaper and faster. An attacker may still need to discover reachable systems, identify the product and protocol, understand the controller configuration, determine which variables or operations matter, and acquire sufficient knowledge of the underlying process. AI can nevertheless accelerate several intermediate tasks: generating or adapting protocol-handling code, modifying existing scripts, interpreting technical documentation, debugging interaction logic, and iterating against industrial libraries and interfaces.
AI therefore changes the economics of capability development more readily than it changes the physics of the target process. Generating a syntactically valid S7 request is fundamentally different from knowing which memory area, operating mode, interlock, setpoint, or control action will produce a desired physical effect. The latter still depends on the architecture, controller configuration, engineering logic, and operating context of the specific installation.
That distinction should shape the defensive response. The relevant problem is not merely that an attacker may possess better scripting assistance. It is that increasingly accessible tooling can shorten the path from identifying an exposed or reachable industrial asset to exercising whatever authority the architecture makes available through it. Removing unnecessary exposure, restricting engineering functions, narrowing controller privileges, monitoring unusual use of legitimate industrial protocols, and independently constraining unsafe physical states therefore become more important as the cost of basic technical interaction falls.
The July 2026 FBI and EPA warning to U.S. water utilities provides complementary evidence. Since July 27, water and wastewater utilities in at least seven states had reported incidents involving internet-facing Allen-Bradley MicroLogix 1100 and 1400 PLCs. Attackers changed IP addresses and passwords, causing loss of monitoring and control; the agencies also described project-file discrepancies and operational degradation, including loss of pressure and flooding.6
Greece’s National Cybersecurity Authority subsequently reported identifying 63 internet-exposed Siemens S7 controllers while responding to the August warning.7 Exposure is not compromise. It is nevertheless architecturally significant because it converts what may have been treated as an internal industrial trust relationship into an externally reachable attack surface.
The deeper architectural lesson appears in Poland. CERT Polska’s investigation of the December 2025 energy-sector attacks showed that direct public-internet exposure of the final industrial target was not necessary. In the most revealing case, compromise of a remote energy installation provided entry into a private cellular APN whose configuration allowed arbitrary endpoints inside the private network to communicate. That path ultimately exposed the OT environment of a separate combined heat and power plant.8
The implication is fundamental: absence of direct internet exposure does not imply isolation. A controller can be unreachable from the public internet yet remain reachable through a compromised remote site, a permissive carrier network, a vendor VPN, an enterprise route, a shared management plane, or another trusted system. In such architectures, the relevant unit of analysis is not the endpoint alone but the complete structure of dependencies and prerequisites through which additional authority can be derived.
The security object is therefore the cyber-physical control path: not merely a network route or a linear sequence, but the causal structure of identities, endpoints, trust relationships, networks, applications, protocols, privileges, and control functions through which an initial digital foothold can acquire process-relevant authority.
This leads to the central thesis of the article:
The primary OT cybersecurity problem is not the existence of vulnerable PLCs in isolation. It is the existence of cyber-physical control paths whose reachability, trust, privilege, integrity, and failure propagation are insufficiently constrained.
A resilient architecture must accordingly be designed on the assumption that individual controls can fail. A stolen credential, compromised firewall, vulnerable controller, trusted engineering workstation, private carrier network, or third-party access channel should not, by itself, provide sufficient authority to produce an uncontrolled physical consequence.
The objective is therefore not perfect prevention. It is bounded compromise. Reachability should be limited to mission-required paths, and each transition toward higher-consequence control should require more specific authority. Privilege should become narrower and more contextual as access approaches the process. High-consequence operations should require stronger and preferably independent controls. Misuse of legitimate authority should be observable. Safety mechanisms should retain meaningful independence from ordinary control. After compromise, the operator should retain a trustworthy path back to a known operational state.
The organizing principle for the analysis that follows is consequently a shift in abstraction: from securing individual devices to constraining the end-to-end accumulation of operational authority.
What the 2026 advisories actually establish
Recent OT advisories should be treated as evidence about different stages of adversary activity, not as interchangeable proof of industrial compromise. Five states must remain analytically distinct:
exposure: a system can be discovered or reached;
targeting: an actor is actively attempting to interact with it;
compromise: unauthorized access, modification, or control has occurred;
operational effect: the compromise has altered, degraded, or interrupted the industrial mission;
attribution: responsibility has been assigned, with some stated level of confidence, to an actor, group, or state.
These states are related, but none logically implies the next. An internet-visible PLC is not necessarily compromised. A compromised controller has not necessarily caused a physical effect. A disruptive incident does not, by itself, establish who caused it.
That distinction is essential for interpreting the 2026 evidence.
The April 2026 joint U.S. advisory AA26-097A represents the strongest evidentiary category considered here because it describes observed exploitation with operational consequences, rather than exposure or vulnerability alone. The authoring agencies assessed that Iranian-affiliated actors were targeting and exploiting internet-connected OT across U.S. critical infrastructure. The affected industrial ecosystems included Rockwell Automation, Schneider Electric, and Siemens, and the activity involved legitimate engineering applications such as Studio 5000, EcoStruxure Control Expert, and TIA Portal.9
The advisory describes manipulation or deletion of PLC project logic and Add-On Instructions, HMI and SCADA manipulation, and interference with alarm or shutdown logic. In one victim environment, a malicious controller project preserved enough expected downstream behavior to avoid immediate detection while introducing logic capable of overriding instructions intended to keep the process within safe parameters. The agencies reported operational disruption and financial loss.10
This is therefore evidence of a complete progression from access to industrial authority and from industrial authority to operational effect.
The August 2026 advisory AA26-231A documents a different point on that progression. It describes targeted reconnaissance and capability development against Siemens S7 environments using Censys, ZoomEye, snap7.dll, python-snap7, and S7comm. The identified tooling was capable of both reading from and writing to controllers, but the publicly described activity emphasized reconnaissance and read operations used to characterize targets and prepare for possible subsequent manipulation.11
The distinction is material. An actor able to enumerate controllers, interrogate their state, understand configurations, and develop write-capable tooling has acquired meaningful technical capability. But such evidence does not demonstrate that the actor has already exercised that capability against every target, still less that it has produced the physical consequences that successful manipulation might permit.
For that reason, the advisory’s discussion of process disruption, equipment damage, safety effects, downtime, and cascading failures should be read as a description of credible consequences of successful exploitation, not as a catalogue of outcomes already observed across the campaign.
The July 2026 FBI and EPA warning to U.S. water and wastewater utilities sits further along the evidentiary chain. Utilities in at least seven states reported unauthorized access to internet-facing Allen-Bradley MicroLogix 1100 and 1400 PLCs. Attackers changed IP addresses and passwords, producing loss of monitoring and control; agencies also reported discrepancies in controller project files and operational effects including pressure loss and flooding.12
Here, exposure had already become compromise, and compromise had already translated into operational degradation. The agencies did not, however, publicly attribute the incidents to a named state actor. The technical evidence is therefore stronger than the attribution evidence.
The Greek National Cybersecurity Authority reported another, more limited evidentiary state: while responding to the August Siemens warning, it identified 63 internet-exposed S7 controllers.13 That finding establishes exposure. It does not establish compromise of those 63 devices.
The United Kingdom’s August 27 NCSC warning operates at a different level again. It describes increased targeting of OT across multiple sectors and jurisdictions, notes that some activity has caused limited real-world disruption, and emphasizes exposure reduction, edge-device security, lifecycle management, segmentation, monitoring, and resilience.14 It is therefore best interpreted as evidence of a broader threat pattern rather than as forensic documentation of one uniform campaign.
Table 1 summarizes the resulting evidentiary hierarchy.
Reconnaissance and capability development observed; read/write-capable tooling described
Potential consequences, not generally demonstrated
No public state attribution
Greek S7 warning, August 2026
63 exposed controllers identified
Wider campaign acknowledged
Not established for those 63 systems
Not established
No Greek attribution
NCSC warning, August 27, 2026
Internet and edge exposure emphasized
Increased targeting reported
General warning rather than case-specific confirmation
Some limited real-world disruption reported
Multiple actors referenced; no attribution of the reported UK generator incident
Table 1: Evidence distinctions in the principal 2026 warnings. Exposure, targeting, compromise, operational effect, and attribution describe different evidentiary states and should not be treated as interchangeable.
This evidentiary discipline is not merely editorial. It changes the defensive problem:
If the evidence establishes only exposure, the immediate objective is to reduce unnecessary reachability and verify the architecture assumed to provide isolation.
If targeting is occurring, monitoring, threat hunting, logging, and scrutiny of engineering and remote-access activity become more urgent.
If compromise is confirmed, the response expands to containment, credential invalidation, integrity verification, forensic preservation, and investigation of every trust relationship through which the attacker may have derived additional authority.
If an operational effect has occurred, digital recovery alone is insufficient: the organization must also establish the physical state of the process, validate controller and protection logic, and determine whether safe operation can resume.
And where attribution remains uncertain, geopolitical interpretation should not substitute for technical incident response.
The advisories nevertheless reveal a structural shift. Industrial scanning services, publicly available protocol implementations, commercial and open-source engineering libraries, legitimate vendor tooling, and AI-assisted code generation increasingly reduce the effort required to move from identifying a reachable industrial asset to interacting with it technically.
That does not collapse the entire attack chain into a single step. Several conditions must still align. An attacker must identify a relevant target, establish a communication path, obtain or derive sufficient authority, understand enough of the industrial semantics to perform a meaningful operation, and reach a system whose action can influence the process.
The escalation is therefore better understood as a set of conditional capability transitions. Discovery may create an opportunity for reachability; reachability may expose a path to operational authority; operational authority may permit an industrial action; and only some industrial actions produce a material physical or operational consequence. None of those transitions is automatic.
Those conditional transitions are precisely where architecture should intervene. A resilient OT system should offer multiple independent opportunities to stop the progression: eliminate unnecessary exposure; constrain identities and privileges; restrict engineering and controller operations; make abnormal use of legitimate authority observable; preserve independent safety boundaries; and ensure that even successful compromise cannot propagate without limit into the physical process.
The significance of the 2026 advisories is therefore not that every exposed PLC is on the verge of destruction. It is that they provide evidence, at different stages of maturity, that adversaries are increasingly traversing the chain from discovering industrial systems toward acquiring operational authority. The relevant defensive question is how many independent constraints remain between the initial foothold and the physical consequence.
From Poland to water utilities: how digital access becomes physical disruption
The phrase OT attack is too coarse to describe the incidents considered here. Digital compromise can affect an industrial mission through materially different mechanisms, and those mechanisms imply different architectural failures.
Three categories are useful:
Direct OT manipulation: the adversary changes a controller state, process setting, actuator command, protection function, or other industrial variable that directly influences physical behavior.
OT impairment: the adversary degrades monitoring, engineering, communications, supervision, or control capability without sufficient evidence that the physical process itself was directly manipulated.
Digitally induced operational disruption: compromise of a supporting digital dependency interrupts the industrial mission even though field-level control equipment is not known to have been manipulated.
The distinction is causal rather than semantic. A plant can lose production because a PLC was deliberately placed in STOP mode, because operators lost supervisory control, or because an indispensable logistics or information system became unavailable. All three are cyber-induced operational events, but they traverse different paths from digital compromise to physical consequence. The December 2025 attacks investigated by CERT Polska illustrate this progression particularly clearly.
Renewable grid-connection points: destructive administration without generation loss
On December 29, 2025, coordinated attacks affected more than 30 Polish wind and photovoltaic installations. The targets included grid-connection environments containing RTUs, HMIs, protection relays, serial servers, routers, firewalls, and other industrial equipment.15
The attackers did not require specialized process malware for each device. CERT Polska documented destructive use of ordinary administrative capabilities, including:
firmware corruption on Hitachi Energy RTU560 devices;
exploitation of default or weak web, SSH, FTP, and local-administrator credentials;
deletion of files from industrial systems;
destructive reset or reconfiguration of communications equipment;
deployment of the DynoWiper wiper against Windows systems.16
The attacks severely degraded telecontrol and supervisory capability, yet electricity generation at the affected renewable installations continued.
That distinction is important. The adversary had compromised systems belonging to the operational environment and had destroyed functions necessary for supervision and remote control, but the physical generation process remained capable of operating. The event is therefore best understood primarily as OT impairment, not as direct manipulation of generation.
It also demonstrates a recurring property of industrial attacks: substantial operational damage does not necessarily require novel ICS malware. Legitimate administrative interfaces, ordinary credentials, and vendor-supported management functions may already expose enough authority to impair the industrial mission.
The large CHP plant: enterprise privilege becomes industrial risk
A separate Polish case involving a large combined heat and power plant exposed a different mechanism.
CERT Polska found that the adversary had maintained long-lived access to the Windows and Active Directory environment, used remote desktop and administrative tools, stolen privileged credentials, accessed domain-controller material, obtained FortiGate configuration data, and attempted broad deployment of DynoWiper. Reverse proxying and lateral-movement tooling were also identified.17
Endpoint-detection tooling prevented the destructive payload from producing its intended full effect. The architectural significance of the case lies elsewhere: nominal segmentation does not necessarily imply independent containment.
A network diagram may show enterprise and OT systems separated by firewalls, VLANs, or routed boundaries. But if the same compromised identities, administration systems, credentials, or management plane can redefine those boundaries, their effective independence is much weaker than their topology suggests.
A VLAN does not constitute an independent security boundary when the attacker controls the infrastructure that defines it. A firewall provides limited containment if its administration depends on an identity plane already compromised by the adversary. Segmentation must therefore be evaluated not only by data-plane topology but also by the control and management authority governing the boundary itself.
The second CHP plant: operational authority assembled across trust domains
CERT Polska’s August 2026 follow-up disclosed an even more instructive attack against a combined heat and power plant serving approximately 50,000 residents. The incident shut down a steam turbine and the process-water treatment system, interrupting cogeneration. Operators restored service quickly enough that customers did not lose heat or electricity.18
The plant was not reached through direct public-internet exposure of its PLCs. Instead, the attack began elsewhere. After compromising a separate wind-farm environment, the adversary reached cellular routing infrastructure associated with that site and used it to enter a private APN connecting industrial installations. The APN permitted arbitrary endpoints within the private network to communicate with one another. Compromise of one connected site therefore provided reachability toward systems belonging to another industrial environment.19
From that position, the attacker identified systems associated with the CHP environment, including a WAGO PFC200 controller and Siemens S7 PLCs. CERT Polska reconstructed reconnaissance over HTTP, VNC, S7, and Modbus. The WAGO controller accepted default administrative credentials. The attacker enabled SSH, established a tunnel through the controller, and used it as a pivot into the CHP OT environment. From there, the adversary interacted with S7-300, S7-1200, and S7-1500 controllers, placed PLCs into STOP mode, and applied password protection. Those actions contributed directly to shutdown of the steam turbine and process-water treatment system.20
The important property of this attack is its composition. No single step was equivalent to control of the CHP process:
flowchart TD
A([Remote-site compromise])
B[Network-path control]
C{Permissive private APN<br/>with insufficient peer isolation}
D[Reachability of another<br/>industrial endpoint]
E[[Administrative authority<br/>via default credentials]]
F[/Tunnel into the<br/>CHP environment/]
G[Reachability of<br/>Siemens PLCs]
H{{Controller operating-mode<br/>authority}}
I(((Process disruption)))
A -->|compromise yields control<br/>of communications path| B
B -->|uses available carrier path| D
C -.->|permits cross-site routing| D
D ==>|default credentials<br/>convert reachability into privilege| E
E -->|privileged authority enables<br/>SSH tunnelling| F
F -->|tunnel extends the<br/>reachable OT domain| G
G ==>|management access exposes<br/>high-consequence functions| H
H ==>|CPU mode change alters<br/>industrial operation| I
Figure 1: Causal progression from compromise of a remote site to physical process disruption. Node shape distinguishes compromise, reachability, trust-boundary weakness, privileged administrative authority, pivot mechanisms, controller authority, and physical consequence; edge labels identify the mechanism by which authority accumulates.
Each transition converted one kind of capability into another. The APN provided reachability, but not PLC control. Default credentials provided administration of the WAGO device, but not yet control of the turbine. The tunnel provided access to another network domain, but physical consequence emerged only after legitimate PLC functionality became available for disruptive use. This is the distinction between initial access and derived operational authority.
A security assessment limited to the final PLC would miss most of the causal structure that made the incident possible. Figure 1 isolates the capability transitions, while Figure 2 expands the same progression across the architectural trust domains in which those transitions occurred.
%%{init: {"theme": "neo", "look": "handDrawn", "layout": "elk"}}%%
flowchart TD
subgraph EXT["External attacker"]
A([Initial attacker])
end
subgraph REMOTE["Compromised remote energy site"]
B([Remote-site compromise])
C[Control of cellular router<br/>or network path]
end
subgraph CARRIER["Private carrier domain"]
D[Private APN<br/>shared carrier domain]
E{Peer isolation absent}
end
subgraph EDGE["CHP-connected industrial edge"]
F[Reachability of WAGO PFC200]
G{Default administrative<br/>credentials retained}
H[[Administrative authority<br/>over WAGO controller]]
I[/SSH tunnel / pivot/]
end
subgraph CHPOT["CHP OT control domain"]
J[Reachability of CHP OT services]
K[Reachability of Siemens S7 PLCs]
L{{Controller operating-mode<br/>authority}}
end
subgraph PROCESS["Physical process"]
M[Steam turbine]
N[Process-water treatment]
O(((Cogeneration interrupted)))
end
A -->|initial intrusion| B
B -->|compromise yields control<br/>of site communications path| C
C -->|access to shared carrier domain| D
D -->|carrier path available| F
E -.->|absence of peer isolation permits<br/>cross-site reachability| F
F ==>|reachable administrative service<br/>can expose privilege| H
G -.->|default credentials satisfy<br/>authentication condition| H
H -->|privileged authority enables<br/>tunnelling| I
I -->|pivot across site boundary| J
J -->|reconnaissance and service discovery| K
K ==>|available functions expose<br/>CPU STOP / password protection| L
L ==>|controller authority changes<br/>physical operation| M
L ==>|controller authority changes<br/>physical operation| N
M -->|loss of process function| O
N -->|loss of process function| O
Figure 2: Authority accumulation in the Polish private-APN incident. Node shapes distinguish compromise states, reachability capabilities, trust-boundary weaknesses, privileged authority, pivot mechanisms, controller authority, and physical consequence. Edge styles distinguish ordinary progression, implicit trust, and authority escalation.
The lesson is not that private APNs are intrinsically insecure. It is that private transport and endpoint isolation are different security properties.
A carrier network can successfully prevent traffic from traversing the public internet while simultaneously allowing every connected endpoint to reach every other endpoint. If the industrial mission does not require that connectivity, the architecture has reduced public exposure without adequately constraining lateral movement.
This distinction is particularly important in distributed OT. Carrier networks, site-to-site VPNs, vendor networks, telemetry systems, and remote-maintenance infrastructure routinely connect facilities that are geographically and operationally distinct. A compromise at one site should not automatically inherit reachability to the others.
The appropriate design question is therefore not:
Is the network private?
It is:
Which endpoint may communicate with which other endpoint, in which direction, through which service, and for what industrial purpose?
Privacy can protect transport. Isolation must be explicitly engineered.
Recent incidents reveal multiple paths from cyber access to operational consequence
The same causal distinction appears across recent international incidents. The cases below are not a statistically representative dataset and cannot support claims about global incident frequency. They are useful because authoritative public reporting provides enough technical information to identify which form of authority or digital dependency was reached before the industrial consequence occurred.
In early 2024, U.S. and international authorities documented attacks against water and wastewater facilities in which pro-Russia hacktivists accessed exposed HMIs, often through VNC protected by default or weak credentials and without multifactor authentication. Attackers altered setpoints and other parameters, disabled alarms, changed administrative passwords, and caused pumps and blower equipment to operate outside expected parameters. Some facilities experienced minor tank overflows, while operators at several sites contained the effects by reverting to manual control.21
These incidents are examples of direct OT manipulation. The attacker did not require exploitation of PLC firmware or purpose-built industrial malware because the legitimate HMI already exposed process-relevant authority.
A December 2025 multinational advisory reported that similar exploitation of internet-visible VNC-connected HMIs and SCADA environments had continued into 2025. The agencies observed that the actors often possessed limited process-specific expertise and sometimes exaggerated their public claims, yet some compromises nevertheless produced real operational effects and physical damage.22
That combination is analytically important: limited sophistication at the attacker side can still produce significant consequences when the compromised interface already provides high-consequence authority.
Denmark provides a clearer physical example. In late 2024, attackers compromised the operational environment of a small water utility and manipulated water pressure. Approximately 450 households temporarily lost supply because of low pressure; subsequent increased pressure caused a pipe to rupture, leaving roughly 50 households without water for several hours.23 In December 2025, Danish Defence Intelligence attributed the destructive attack to the pro-Russian group Z-Pentest and assessed that the group had connections to the Russian state.24
Here, digital access became direct authority over a physical process variable, and manipulation of that variable propagated into infrastructure damage.
The April 2025 Bremanger dam incident in Norway follows the same causal pattern in a different domain. Norwegian authorities reported compromise of the dam’s control system and modification of operating settings that left a water-control mechanism open for approximately four hours, releasing about 500 litres of water per second. The overall damage was limited, but the causal mechanism was direct: unauthorized access to an industrial control interface became authority over a physical actuator.25
North Carolina Ports in August 2026 illustrates a fundamentally different path. A cyberattack disrupted information systems across the Port of Wilmington, the Port of Morehead City, and the Charlotte Inland Port. The resulting outage delayed gate operations and degraded terminal activity while contingency procedures and restoration were underway.26
Public reporting does not establish compromise of cranes, PLCs, or other field-level control equipment. The incident is therefore better classified as digitally induced operational disruption. The industrial mission was impaired because it depended on digital systems that had become unavailable, not because attackers were publicly shown to have manipulated the physical-control layer.
The Polish cases occupy two different positions in this taxonomy. At renewable grid-connection points, destructive administration of RTUs, HMIs, communications equipment, and supporting Windows systems caused loss of telecontrol while generation continued: principally OT impairment.27 At the second CHP plant, compromise propagated through a remote site, carrier network, industrial controller, and PLC environment until the attacker acquired operating-mode authority and contributed directly to process shutdown: direct OT manipulation through composed authority.28
The July 2026 U.S. water-sector incidents show another mixed case. The FBI and EPA reported unauthorized access to internet-facing Allen-Bradley MicroLogix 1100 and 1400 PLCs. Attackers modified IP addresses and passwords, causing loss of monitoring and control, while at least one affected organization detected discrepancies in PLC project logic. Reported consequences included pressure loss and flooding.29
Because the compromised controllers performed different operational functions, the incidents span both OT impairment and direct controller manipulation with operational consequence rather than forming one uniform technical category.
Table 2 compares the cases according to the capability or dependency that was actually reached.
Case
Initial access or dependency failure
Authority or dependency reached
Observed consequence
Analytical classification
U.S. water and wastewater facilities, early 2024
Internet-exposed VNC/HMI; weak or default credentials; absent MFA
HMI setpoints, alarms, settings, and administrative functions
Pumps and blowers outside normal parameters; minor tank overflows at some sites; manual fallback used
Direct OT manipulation
Danish water utility, late 2024
Weakly protected OT environment
Water-pressure control
Loss of pressure, subsequent pressure increase, pipe rupture, temporary interruption of supply
Direct OT manipulation
Bremanger dam, Norway, April 2025
Compromise of remotely accessible control system
Water-control operating settings
Approximately 500 L/s released for about four hours
Direct OT manipulation
Polish renewable sites, December 2025
Perimeter compromise, weak credentials, destructive use of legitimate management functions
RTUs, HMIs, protection and communications infrastructure
Steam-turbine and process-water shutdown; cogeneration interrupted
Direct OT manipulation through composed authority
U.S. water utilities, July 2026
Internet-facing PLC access
PLC network configuration and, at some sites, controller project state
Loss of monitoring/control, pressure loss, flooding
Mixed OT impairment and direct manipulation
North Carolina Ports, August 2026
Cyberattack against supporting information systems
Digital dependencies required for gate and terminal operations
Systems outage, gate delays, degraded terminal operation
Digitally induced operational disruption
Table 2: Recent incidents illustrate different causal mechanisms by which cyber compromise becomes industrial consequence. The taxonomy is analytical rather than an official incident classification.
The comparison supports a more useful conclusion than the proposition that OT attacks are simply becoming more sophisticated.
Several consequential incidents relied on techniques that are technically ordinary: exposed VNC services, default credentials, legitimate HMIs, standard controller functions, remote-access infrastructure, shared carrier networks, and dependence on centralized digital services.
What determines consequence is not primarily the novelty of the exploit. It is the authority or dependency exposed by the compromised path. An attacker who compromises a read-only historian has acquired a different capability from one who compromises an HMI capable of changing pressure. Membership in a private APN is different from PLC programming authority, but it may become the first element in a sequence that ultimately derives that authority. Destruction of an enterprise application is different from manipulation of field control, yet it can still stop the industrial mission when operations depend on that application.
The relevant causal quantity is therefore not simply access, but the progressive accumulation of operational authority. An initial digital foothold becomes consequential only when the surrounding architecture allows it to acquire additional reachability, exploit trust-boundary weaknesses, obtain privileged administrative capabilities, and ultimately exercise authority over the industrial process.
The incidents differ in where this progression begins, which architectural dependencies permit it to continue, and at which transition ordinary digital compromise becomes high-consequence operational authority.
In some cases, the first compromised interface already contains substantial process authority. In others, authority is accumulated through several intermediate systems. In still others, no process-control authority is required because the industrial mission depends critically on a digital service whose loss is sufficient to stop operations.
This distinction has a direct architectural consequence: security controls should be placed according to the authority and dependency structure of the system, not merely according to network location or product type. The recurring variable across these incidents is therefore neither the malware family nor the PLC manufacturer. It is what the compromised path ultimately allows the adversary to control, disable, redefine, or prevent, and what physical or operational consequence lies beyond that capability.
A taxonomy of operational technology cyber risk
The incidents examined above suggest that vulnerable PLC is not an adequate unit of analysis for OT cybersecurity. The relevant failures occur at different points in the cyber-physical system. Some create unnecessary reachability. Some allow reachability to become authority. Some undermine the integrity of the systems through which authority is exercised. Some prevent defenders from determining what has happened. Others determine whether successful compromise remains bounded or propagates into prolonged operational or physical consequence.
A useful taxonomy must preserve those distinctions.
The model used here identifies sixteen failure classes, organized into six broader risk domains:
knowledge and governance: whether the organization understands the cyber-physical system and owns its high-consequence relationships;
reachability and trust: whether communication paths and trust relationships are limited to those required by the operational mission;
identity and authority: whether actors can be identified and whether their effective privileges correspond to their legitimate role;
control-path integrity: whether the systems that create, transmit, and execute process-relevant decisions remain trustworthy;
observability: whether misuse of authority and divergence between digital and physical state can be detected and interpreted;
safety and resilience: whether successful compromise remains bounded and whether trustworthy operation can be maintained or reconstructed.
These are not sequential defensive layers, nor are the sixteen classes mutually exclusive. A single incident can instantiate several simultaneously. The Polish private-APN attack, for example, combined unnecessary reachability, a weak trust boundary, shared infrastructure, weak administrative identity, excessive downstream authority, controller exposure, insufficient containment, and ultimately direct access to process-relevant functions.
The purpose of the taxonomy is precisely to prevent those mechanisms from disappearing into a generic label such as PLC compromise.
%%{init: {"theme": "neo", "look": "handDrawn", "layout": "elk"}}%%
flowchart LR
subgraph KG["Knowledge and governance"]
AK{Asset-knowledge<br/>failure}
OO{Organizational-ownership<br/>failure}
end
subgraph RT["Reachability and trust"]
EX{Exposure<br/>failure}
TB{Trust-boundary<br/>failure}
SI{Shared-infrastructure<br/>failure}
end
subgraph IA["Identity and authority"]
ID{Identity<br/>failure}
PR{Privilege<br/>failure}
TP{Third-party-authority<br/>failure}
end
subgraph CI["Control-path integrity"]
PS{Protocol-semantic<br/>failure}
EW{Engineering-workstation<br/>failure}
CT{Controller-integrity<br/>failure}
end
subgraph OB["Observability"]
PO{Process-observability<br/>failure}
MO{Monitoring<br/>failure}
end
subgraph SR["Safety and resilience"]
SB{Safety-boundary<br/>failure}
AD{Availability-dependency<br/>failure}
RI{Recovery-integrity<br/>failure}
end
R1[Unknown or unnecessary<br/>reachability]
R2[[Unaccountable or excessive<br/>operational authority]]
R3{{Untrusted or unverifiable<br/>controller / process state}}
R4[Reduced ability to detect<br/>or interpret compromise]
R5(((Unsafe, unavailable, or<br/>difficult-to-recover operation)))
%% Failure classes enable architectural effects
AK -.->|incomplete architecture knowledge| R1
EX -.->|unnecessary exposure| R1
TB -.->|boundary does not constrain trust| R1
SI -.->|shared infrastructure propagates reachability| R1
ID -.->|identity cannot be strongly established| R2
PR -.->|privilege exceeds mission need| R2
TP -.->|external authority is insufficiently bounded| R2
PS -.->|permitted protocol access exposes dangerous operations| R2
EW -.->|trusted engineering path is compromised| R3
CT -.->|deployed controller state cannot be trusted| R3
PO -.->|physical state cannot be independently established| R3
PO -.->|loss of independent process evidence| R4
MO -.->|security-relevant activity becomes less visible| R4
SB -.->|common cyber failure reaches protection| R5
AD -.->|dependency loss affects mission| R5
RI -.->|trusted state cannot be reconstructed| R5
SI -.->|common-mode dependency amplifies consequence| R5
OO -.->|unowned risk persists or recovery is delayed| R5
%% Principal authority and consequence escalation
R1 ==>|reachability can be converted<br/>into privilege| R2
R2 ==>|privileged authority can modify<br/>controller or engineering state| R3
R3 -.->|compromise can suppress or<br/>corrupt evidence| R4
R3 ==>|untrusted control state can<br/>alter the physical mission| R5
R4 -.->|delayed detection or response<br/>can increase consequence| R5
Figure 3: Taxonomy of OT cyber risk. Six domains contain sixteen failure classes. Diamond-shaped nodes represent architectural weaknesses or failure conditions. The lower nodes represent their principal effects: unnecessary reachability, excessive operational authority, untrusted or unverifiable controller or process state, impaired observability, and physical or mission consequence. Dashed edges represent contributory relationships; thick edges represent escalation toward higher-consequence authority or effect.
The figure is a map of risk composition, not a universal attack sequence. Availability-dependency failure can stop an industrial mission without controller compromise. Process-observability failure can occur without an external attacker. Organizational-ownership failure does not itself create an exploit, but it can leave an important end-to-end dependency unmanaged. Conversely, a single well-placed control can prevent an otherwise serious compromise from acquiring physical significance.
Knowledge and governance
Knowledge and governance failures determine whether the organization understands the system whose risk it is attempting to control.
Asset-knowledge failure
Asset-knowledge failure occurs when the organization lacks an authoritative representation of the devices, communication relationships, identities, dependencies, external services, and third-party connections capable of affecting the industrial mission.
A conventional inventory of device names, manufacturers, and IP addresses is therefore insufficient.
Security decisions depend on relationships:
which engineering workstation can program which controller;
which HMI can issue commands to which process cells;
which remote sites share carrier infrastructure;
which supplier identities can initiate remote maintenance;
which controllers depend on a common gateway;
which boundaries depend on a common identity or management plane;
which industrial functions depend on cloud, licensing, telecom, or enterprise services;
which applications can ultimately issue process-relevant commands.
NCSC guidance on maintaining a definitive view of OT architecture similarly emphasizes dependencies, connectivity, data flows, and third-party relationships rather than asset enumeration alone.30
The distinction is fundamental. An unknown asset is a visibility problem. An unknown relationship can be a capability-derivation problem. An undocumented vendor route cannot be intentionally constrained. An unknown APN relationship cannot be tested for peer isolation. An unidentified shared identity provider cannot be assessed as a common-mode dependency. A controller whose authorized engineering peers are unknown cannot be monitored effectively for abnormal programming activity.
The required property is therefore not merely inventory completeness. It is architectural knowledge sufficient to explain how reachability, authority, dependency, and consequence propagate through the system.
Organizational-ownership failure
Organizational-ownership failure occurs when individual components are assigned to responsible parties but no person or function owns the complete high-consequence relationship formed by their composition. OT architectures routinely cross organizational boundaries:
industrial networking;
automation engineering;
cybersecurity;
operations;
functional safety;
telecommunications;
procurement;
system integration;
equipment vendors;
cloud and managed-service providers.
Each participant can make a locally defensible decision while the resulting composition remains unsafe. A telecommunications team may classify a private APN as secure connectivity. An automation team may correctly configure the PLC. A supplier may regard persistent remote access as operationally convenient. Cybersecurity may consider a perimeter firewall the OT boundary. Yet those decisions can compose into a path that allows compromise of one site to acquire programming authority at another.
The required property is therefore end-to-end ownership of high-consequence derivations. For every high-consequence cyber-physical relationship, someone must be accountable for explaining why it exists, which prerequisites and assumptions it depends on, which controls constrain it, how those controls are verified, and how the residual risk is accepted.
Reachability and trust
Reachability and trust failures determine where compromise can propagate after the initial foothold.
Exposure failure
Exposure failure occurs when an asset, service, or management interface is reachable from a source that has no operational requirement to communicate with it. Public-internet exposure is only the most visible case. Unnecessary reachability can also arise through:
enterprise routing;
site-to-site VPNs;
private APNs;
carrier networks;
vendor remote-access infrastructure;
shared management systems;
cloud connectivity;
poorly constrained inter-site routing.
The relevant question is therefore not:
Is this PLC internet-facing?
It is:
Which principals and systems can establish a path to it, through which conduits, and why does the industrial mission require each of those paths?
Connectivity inherited from routing convenience, historical topology, supplier defaults, or an assumption that a network is trusted should be treated as exposure until its operational necessity is demonstrated. Exposure reduction consequently means removing all unnecessary reachability, not merely public IP addresses.
Trust-boundary failure
Trust-boundary failure occurs when a component appears to separate security domains but does not remain an effective constraint under the compromise scenarios it is intended to contain.
A firewall may separate enterprise and OT networks while its administration depends on enterprise credentials already controlled by the attacker. A private WAN may remove traffic from the public internet while permitting unrestricted communication among participating industrial sites. A jump server may technically mediate remote access while retaining persistent credentials that provide broad downstream authority.
In each example the boundary exists topologically, but its security independence is weaker than the diagram suggests. The relevant question is therefore not:
Is there a firewall, VLAN, VPN, DMZ, or private network?
It is:
Which attacker capabilities does that boundary continue to constrain after an adjacent system, identity, endpoint, or management plane has been compromised?
A meaningful trust boundary should independently restrict one or more properties such as source, destination, identity, service, protocol operation, administrative authority, or direction of information flow. The greater the consequence beyond the boundary, the stronger the case for ensuring that its enforcement cannot be silently redefined by the environment it is intended to contain.
Shared-infrastructure failure
Shared-infrastructure failure occurs when nominally independent sites, assets, or process cells depend on a common technical component whose compromise can affect several of them simultaneously.
Examples include:
one APN connecting many remote plants;
one identity provider controlling several OT environments;
one firewall-management system defining multiple boundaries;
one engineering repository serving an entire fleet;
one cloud platform required by geographically distributed assets;
one supplier identity accepted across many installations.
The essential property is correlated failure.
Centralized infrastructure is not intrinsically undesirable. Central identity, monitoring, configuration management, or communications can improve consistency and security. The risk arises when the architecture treats a component as merely supportive even though compromise of that component can redefine trust or authority across several otherwise independent environments.
The relevant security question is therefore whether a shared dependency constitutes a common propagation mechanism or common point of authority.
Identity and authority
Identity and authority failures determine whether reachability can be transformed into operational authority.
Identity failure
Identity failure occurs when the architecture cannot reliably establish which human, device, application, or service is exercising a capability. Typical mechanisms include:
default passwords;
shared controller accounts;
generic service accounts;
permanent supplier credentials;
anonymous protocol access;
unauthenticated engineering services;
weak shared communities or secrets.
The problem extends beyond authentication strength. If several actors share one credential, a technically authenticated action may still be unaccountable. The system cannot distinguish which operator, endpoint, application, supplier session, or maintenance activity generated the command.
Without attributable identity, precise authorization and meaningful forensic interpretation become difficult or impossible. The required property is therefore unique and accountable identity wherever the technology permits it, with compensating architectural controls where legacy technology does not.
Privilege failure
Identity answers:
Who or what is acting?
Privilege answers:
Which industrial operations may that actor perform?
Privilege failure occurs when an authenticated principal can exercise greater operational authority than its legitimate function requires. An HMI operator may need to start and stop equipment but not download controller logic. A historian may require process reads but no write authority. A vendor engineer may require temporary programming access to one controller without persistent access to an entire plant.
The relevant quantity is therefore not generic access, but effective operational authority: the set of process-relevant actions technically available to a principal compared with the set required by its mission.
Privilege should generally narrow as access approaches the physical process. Firmware modification, project download, controller mode change, credential administration, protection-setting modification, and unrestricted process write authority should be exceptional rather than incidental consequences of network access.
Third-party-authority failure
Third-party-authority failure occurs when a supplier, integrator, maintainer, carrier, or managed-service provider possesses technical authority broader, more persistent, less revocable, or less observable than its operational role requires.
Supplier risk is therefore not reducible to the supplier’s own cybersecurity maturity. The architecture must examine delegated authority.
For each third party, the operator should be able to determine:
which assets can be reached;
through which identities and endpoints;
whether access is permanent or activated on demand;
which operations become available after access is granted;
whether the operator can independently revoke access;
whether sessions and engineering actions are recorded;
whether compromise of the supplier exposes one asset, one site, or an entire fleet.
The desired property is authority bounded in scope, duration, target, and operation, with activation and revocation remaining under the asset owner’s control.
Control-path integrity
These failures concern the mechanisms through which digital decisions acquire physical meaning.
Protocol-semantic failure
Protocol-semantic failure occurs when the architecture constrains transport connectivity but not the industrial meaning of the operations carried over that connectivity. A firewall rule allowing TCP/102 may constrain which systems can establish Siemens S7 communication. It does not necessarily distinguish among diagnostic reads, project transfer, parameter modification, programming operations, or controller-mode changes.
This distinction extends across industrial protocols. Network authorization asks:
May A communicate with B?
Operational authorization asks:
Which principal, through which trusted application, may perform which industrial operation against which target and under which operating condition?
The second question is much closer to the actual cyber-physical risk. Modern systems may provide protocol- or controller-level mechanisms for enforcing parts of that policy. Legacy environments often cannot. Where semantic authorization is unavailable in the protocol itself, the missing constraint must be provided through architecture: dedicated engineering stations, strict conduits, role separation, maintenance-state controls, application proxies, controller protections, or physical and procedural restrictions.
Engineering-workstation failure
Engineering-workstation failure occurs when a system containing legitimate programming tools, controller projects, credentials, certificates, and trusted communication paths is treated as an ordinary endpoint.
An engineering workstation frequently contains exactly the capabilities an adversary would otherwise need to construct:
vendor programming environments;
authenticated controller sessions;
firmware-management functions;
approved project files;
privileged credentials;
trusted routes through OT boundaries.
Compromise of that system can therefore convert an ordinary workstation compromise directly into process-relevant authority without exploiting the controller itself.
The engineering workstation is consequently part of the controller’s practical trusted computing base. If the engineering environment cannot be trusted, the organization cannot reliably infer that a syntactically legitimate controller modification is an authorized one.
Controller-integrity failure
Controller-integrity failure occurs when the organization cannot establish whether the running controller corresponds to an independently approved state.
That state includes more than application logic. Depending on the device, it may include:
firmware;
controller project;
hardware configuration;
communications configuration;
protection parameters;
operating mode;
local users and credentials;
certificates;
safety-relevant configuration.
The essential question is:
Can the current deployed state be compared with an independently trusted baseline, and can unexplained divergence be detected?
This is different from backup. A backup establishes that another copy exists. Integrity verification establishes whether the running system corresponds to the state the organization has approved.
Observability
Observability failures determine whether defenders can distinguish legitimate operation, accidental failure, adversarial manipulation, and uncertainty.
Process-observability failure
Process-observability failure occurs when the organization cannot establish whether the digital representation of the process corresponds sufficiently closely to the physical process itself.
An HMI can display a plausible tank level while the underlying measurement has been manipulated. A historian can faithfully record false values supplied by a compromised controller. An alarm system can remain quiet because the logic responsible for generating alarms has itself been modified. Multiple displays are not independent evidence when they all depend on the same sensor, controller, or communications path.
Depending on the process, independent evidence can come from:
redundant or diverse sensors;
physical invariants;
command-response consistency;
independent protection systems;
process-model residuals;
controller-state evidence;
local operator observation.
The objective is not necessarily continuous reconstruction of the complete physical state. It is to make important cyber-physical inconsistencies observable.
Monitoring failure
Monitoring failure occurs when the exercise of process-relevant authority either generates insufficient trustworthy evidence or is not interpreted in its operational context.
High-value events include:
controller programming;
project upload or download;
firmware modification;
CPU mode change;
user or credential modification;
protection-setting changes;
unexpected engineering peers;
activation of remote access;
creation of tunnels;
unusual inter-zone communication;
break-glass access;
process behavior inconsistent with command history.
Generic malware detection is insufficient because OT compromise frequently abuses valid protocols, legitimate applications, authenticated identities, and vendor-supported controller functions. The more useful question is therefore:
Did some principal exercise industrial authority in a way inconsistent with its identity, role, endpoint, target, time, approved activity, or physical-process context?
Safety and resilience
These failures determine whether successful cyber compromise remains bounded and whether trustworthy operation survives or can be restored.
Safety-boundary failure
Safety-boundary failure occurs when compromise of ordinary process control can also defeat the mechanisms intended to prevent dangerous physical states.
A Safety Instrumented System or other protective function may be functionally separated from the Basic Process Control System while still sharing:
engineering workstations;
credentials;
identity infrastructure;
network equipment;
maintenance channels;
administrative personnel.
Such dependencies can invalidate the independence implicitly assumed by the safety architecture. The relevant cybersecurity property is therefore not merely security of the safety controller. It is preservation of protective independence under plausible compromise of ordinary control.
Availability-dependency failure
Availability-dependency failure occurs when loss of a supporting digital service prevents the industrial mission from continuing even though field-level control has not been directly manipulated.
Potential dependencies include:
supervisory systems;
production-management platforms;
identity services;
databases;
cloud platforms;
remote licensing;
telecommunications;
scheduling systems;
logistics applications.
The relevant question is:
Which physical functions disappear when this digital dependency disappears?
For each important dependency, the architecture should establish which process functions can continue autonomously, which degrade safely, which can be substituted manually, and which stop entirely. This class explains why a cyberattack against an IT or supervisory dependency can produce industrial disruption without becoming a direct PLC attack.
Recovery-integrity failure
Recovery-integrity failure occurs when services can be restored but the organization cannot establish that the resulting cyber-physical system is trustworthy. Restoring a controller project is insufficient if:
compromised firmware remains installed;
attacker persistence survives elsewhere in the network;
stolen credentials remain usable;
the wrong project revision is restored;
protection settings are unknown;
engineering systems remain compromised;
safety functions have not been independently verified;
the physical process is no longer in the state assumed by the restored automation.
Recovery in OT therefore requires reconstitution of trust, not merely restoration of availability. Normal automated operation should resume only after sufficient evidence exists about identity, network boundaries, engineering systems, controller state, supervisory systems, protection functions, and the corresponding physical process state.
Table 3 condenses the taxonomy into the security question associated with each failure class.
Risk domain
Failure class
Core security question
Knowledge and governance
Asset knowledge
Do we know the assets, relationships, dependencies, identities, and external connections capable of affecting the industrial mission?
Knowledge and governance
Organizational ownership
Who owns the complete high-consequence derivation rather than only its individual components and dependencies?
Reachability and trust
Exposure
Which systems or principals can reach this asset, and which paths are operationally necessary?
Reachability and trust
Trust boundary
Does the boundary continue to constrain compromise after failure of the adjacent environment or its management plane?
Reachability and trust
Shared infrastructure
Can compromise of one common dependency expand authority or consequence across otherwise separate assets or sites?
Identity and authority
Identity
Can the system establish which human, device, application, or service is acting?
Identity and authority
Privilege
Which industrial operations can that principal perform, and are all of them required?
Identity and authority
Third-party authority
Is delegated authority narrow, temporary, revocable, attributable, and observable?
Control-path integrity
Protocol semantics
Does authorization constrain the industrial operation or merely permit network communication?
Control-path integrity
Engineering workstation
Is the engineering environment protected as part of the controller’s trusted computing base?
Control-path integrity
Controller integrity
Can deployed firmware, logic, configuration, mode, and protection state be verified against an approved baseline?
Observability
Process observability
Can actual physical state be distinguished from a compromised digital representation?
Observability
Monitoring
Would abnormal exercise of legitimate industrial authority be visible and interpretable?
Safety and resilience
Safety boundary
Does independent protection remain meaningfully independent under cyber compromise?
Safety and resilience
Availability dependency
Which industrial functions fail when supporting digital services disappear?
Safety and resilience
Recovery integrity
Can the system return to a demonstrably trustworthy digital and physical state?
Table 3: Six risk domains and sixteen distinct OT cybersecurity failure classes. The classes are separated because they alter different properties of the cyber-physical system and therefore require different architectural responses.
From attack surface to derived operational authority
The taxonomy exposes a more fundamental distinction between attack surface and derived operational authority.
Attack surface describes the set of footholds from which an adversary can plausibly enter the system. That is necessary information, but it is not sufficient to characterize industrial cyber risk. Two exposed assets can contribute equally to an attack-surface inventory while having radically different consequences after compromise. A read-only telemetry endpoint may provide observation and little else. An exposed engineering workstation may provide a route into authenticated engineering services, access to controller projects and credentials, and ultimately the ability to modify the physical process. The relevant question is therefore not only:
What can the attacker initially reach?
It is:
What additional capabilities can the attacker derive from that foothold through the trust, identity, network, engineering, and control relationships already present in the architecture?
A capability-inference model
Let V be a finite set of security-relevant attacker capabilities and consequence states. Elements of V can represent capabilities such as:
reachability of a particular service;
possession of a credential;
authenticated access to an engineering environment;
administrative control of a gateway;
ability to program a controller;
ability to issue a process command;
as well as consequence states such as loss of a process function.
A simple directed graph is not always sufficient because acquisition of one capability can require several prerequisites simultaneously. Programming a PLC, for example, may require both network reachability and possession of an appropriate credential. Represent the architecture instead as a set of typed inference rules \mathcal{R}. Each rule r=\left(P_r,v_r,\tau_r\right) contains:
a prerequisite set P_r \subseteq V;
a capability or consequence v_r \in V that can be derived;
a transition class \tau_r \in \mathcal{T}.
The transition classes can include \mathcal{T}=\left\{\mathrm{reach},\mathrm{authenticate},\mathrm{administer},\mathrm{program},\mathrm{command},\mathrm{affect}\right\}. The transition type is descriptive: it records what kind of capability or consequence transition the rule represents. It does not imply that every attack must traverse these classes in one fixed sequence.
Let \mathcal{Z} denote the set of architectural and operational states, and let z \in \mathcal{Z} be the state being analyzed. It includes relevant assumptions such as routing, enabled services, authentication configuration, controller configuration, access policy, and trust relationships. For each rule r, define the architectural-feasibility predicate q_r:\mathcal{Z}\rightarrow\{0,1\}, where q_r(z)=1 means that the architectural conditions required by rule r are present under state z. Examples include:
a route exists;
an administrative service is enabled;
the authentication configuration accepts a credential already available to the attacker;
a conduit permits the relevant traffic;
the controller accepts programming;
a maintenance service is reachable.
The separation between P_r and q_r(z) is useful. P_r represents capabilities the attacker must already possess. q_r(z) represents properties of the deployed architecture that make the transition possible.
Capability closure
Let C \subseteq V be the capabilities already available to the attacker. Define the capability-expansion operator
\Phi_z(C)
=
C
\cup
\left\{
v_r
\;\middle|\;
r \in \mathcal{R},
\;
P_r \subseteq C,
\;
q_r(z)=1
\right\}.
\tag{1}
The operator adds every capability whose prerequisites are already available and whose architectural conditions are satisfied.
Let C_0 \subseteq V represent the capabilities obtained from a plausible initial compromise. Define C_z^{(0)}(C_0)=C_0, and iteratively
Because \Phi_z never removes an acquired capability, C_z^{(k)}(C_0)\subseteq C_z^{(k+1)}(C_0). Since V is finite, this sequence reaches a fixed point after finitely many iterations. Define that fixed point as C_z^{*}(C_0). It satisfies
The set C_z^{*}(C_0) is the capability closure of the initial foothold. It contains every capability or consequence that can be derived by repeatedly applying the trust and authority relationships encoded in the architecture. This is the first useful consequence of the model:
An initial compromise is characterized not only by the capability it immediately provides, but by the closure it generates.
Attack surface and reachable authority are different objects
Let \mathfrak{C}_0=\left\{C_0^{(1)},C_0^{(2)},\ldots,C_0^{(m)}\right\} represent the plausible initial-compromise scenarios exposed by the system’s attack surface. Attack-surface analysis identifies the elements of \mathfrak{C}_0. The closure model evaluates what each of those footholds can become under the deployed architecture.
Now define \mathcal{A} \subseteq V as the subset of capabilities that confer operational or administrative authority. For foothold C_0, define its reachable authority set as
This is the authority reachable from that foothold under architecture z. The authority created by architectural composition, rather than already present in the initial compromise, is
This quantity has a direct security interpretation. If \Delta \mathcal{A}_z(C_0)=\varnothing, then the architecture does not amplify the foothold into additional operational authority.
If \Delta \mathcal{A}_z(C_0)\neq\varnothing, then the attacker can derive authority that was not contained in the original compromise. This is the distinction that an attack-surface count cannot express. For example, suppose C_0^{(\mathrm{telemetry})} represents compromise of one exposed telemetry endpoint and C_0^{(\mathrm{engineering})} represents compromise of one exposed engineering endpoint. An inventory might count both as one exposed asset. But it is entirely possible that \Delta \mathcal{A}_z\left(C_0^{(\mathrm{telemetry})}\right)=\varnothing while \Delta\mathcal{A}_z\left(C_0^{(\mathrm{engineering})}\right)\neq\varnothing. The two exposures are therefore not equivalent even though their contribution to a simple attack-surface count is identical.
High-consequence derivability
The same model connects derived authority to physical consequence. Let H \subseteq V be the set of high-consequence capabilities or states that the architecture is intended to prevent. Examples can include:
unauthorized controller programming;
loss of an independent protection function;
forced controller STOP;
uncontrolled process manipulation;
loss of a safety-critical process function.
A high-consequence state is derivable from foothold C_0 when C_z^{*}(C_0)\cap H\neq\varnothing. Conversely, C_z^{*}(C_0)\cap H=\varnothing means that none of the defined high-consequence states is derivable from that foothold under the modeled architecture.
This yields a second useful consequence of the model. Two systems can expose the same number of externally reachable assets while having very different cyber-physical risk because their capability closures differ. Likewise, removing one externally exposed asset is not necessarily more important than breaking one internal trust relationship if that relationship lies on every inference chain through which several footholds acquire high-consequence authority. The security object is therefore not the exposed asset in isolation. It is the closure generated by compromising it.
The model also clarifies what an architectural control actually does. A firewall rule, removal of a default credential, controller hardening, a dedicated engineering workstation, stronger privilege separation, or peer isolation need not make the initial foothold impossible. Instead, it can make one or more inference rules unavailable by changing q_r(z) from 1 to 0.
The result is a different capability closure. For a purely restrictive control, one that only makes previously feasible inference rules infeasible and introduces no new attacker capabilities or inference rules, the post-control closure satisfies C_{z'}^{*}(C_0)\subseteq C_z^{*}(C_0). A strict inclusion C_{z'}^{*}(C_0)\subsetneq C_z^{*}(C_0) means that at least one capability previously derivable from the same foothold has been removed.
For a general architectural change, however, the two closures need not be nested because a new control can also introduce dependencies, services, identities, or management paths. Its security effect should therefore be evaluated against the high-consequence set directly. For controls intended to prevent high-consequence escalation, the design objective is C_{z'}^{*}(C_0)\cap H=\varnothing.
The value of the control is therefore not merely that it exists. Its value is that it changes the inference system so that particular capabilities can no longer be derived. This gives formal meaning to several architectural controls discussed throughout the article:
segmentation can remove derivable reachability;
identity controls can remove derivable authenticated authority;
privilege controls can remove derivable administrative or engineering authority;
controller hardening can remove derivable programming or command capabilities;
safety independence can prevent ordinary control compromise from deriving protection-system compromise.
The Polish incident as an instance of the model
The Polish private-APN incident illustrates the model without requiring the incident sequence itself to be written as mathematics.
The initial compromise did not directly provide authority over the CHP process. Instead, acquired attacker capabilities combined with enabling architectural conditions to make additional capabilities derivable: control of a communications path combined with permissive cross-site routing, reachability of the industrial controller combined with weak authentication, administrative authority enabled tunnelling into the CHP OT environment, and the resulting reachability of Siemens PLCs exposed high-consequence controller functions.
Figure 1 and Figure 2 represent that causal progression explicitly. In the formal model, the relevant result is that the initial foothold C_0 generated a closure satisfying C_z^{*}(C_0)\cap H\neq\varnothing. The incident therefore demonstrates something more specific than the attacker gained access. It demonstrates architectural amplification: a comparatively remote initial compromise generated capabilities that were not present in the foothold itself and ultimately made a high-consequence industrial state derivable.
That is the distinction between attack surface and reachable authority:
Attack surface identifies where compromise can begin.
Capability-inference analysis determines what the architecture allows that compromise to become.
The model is deliberately non-probabilistic. It does not estimate the probability of the initial compromise, the probability that an adversary will select a particular inference chain, or the probability that every technically feasible action will succeed in practice. It answers a narrower architectural question:
Given this initial foothold and these architectural assumptions, what operational authority or consequence is derivable?
From capability derivations to control placement
The defensive problem follows directly from the capability-closure model.
Let \mathfrak{C}_0=\left\{C_0^{(1)},C_0^{(2)},\ldots,C_0^{(m)}\right\} denote the set of plausible initial-compromise scenarios that the architecture is explicitly designed to tolerate. Let \mathcal{D}=\left\{d_1,d_2,\ldots,d_r\right\} denote the candidate defensive controls. A defensive control can change the inference system in several ways. It may:
remove a network path required by an inference rule;
make an authentication condition false;
reduce the authority available to an authenticated principal;
disable an engineering or programming service;
isolate network peers;
constrain a maintenance relationship;
prevent a system from being used as a tunnel or pivot;
remove a shared dependency;
preserve an independent safety or protection boundary.
These controls operate at different technical layers. In the formal model, however, their relevant effect is the same: they change which attacker capabilities remain derivable from a given foothold.
Controls transform the inference system
For a selected control set K\subseteq\mathcal{D}, let \mathcal{R}_K denote the inference-rule set after those controls are applied. For each rule r \in \mathcal{R}_K, let q_r^K(z)\in\{0,1\} denote its architectural-feasibility predicate under control set K. This notation allows a control to do more than simply delete an existing relationship.
A control can:
make an existing inference rule infeasible;
modify the prerequisites for exercising a capability;
replace one inference rule with a more restrictive one;
introduce a new service or dependency that must itself be represented in the model.
The last case is important. A remote-access broker, identity service, security gateway, or centralized management platform may eliminate dangerous attacker paths while simultaneously introducing new dependencies. Control evaluation must therefore consider the resulting inference system, rather than merely count the transitions that were removed.
Define the post-control capability-expansion operator as
\Phi_{z,K}(C)
=
C
\cup
\left\{
v_r
\;\middle|\;
r \in \mathcal{R}_K,
\;
P_r \subseteq C,
\;
q_r^K(z)=1
\right\}.
\tag{6}
Starting from initial foothold C_0, repeated application of \Phi_{z,K} produces the post-control capability closure C_{z,K}^{*}(C_0), which contains every capability derivable under architectural state z after deployment of control set K.
As in the baseline model, the closure is the fixed point of the post-control expansion operator:
Equation Equation 7 states that, once the closure has been reached, no further capability can be derived by another application of the post-control inference rules. The defensive effect of K can therefore be evaluated directly by comparing which capabilities remain derivable in C_{z,K}^{*}(C_0) with those derivable in the corresponding baseline closure.
High-consequence separation
Let H\subseteq V be the high-consequence capability and state set defined previously. The fundamental architectural requirement is
None of the initial compromises that the architecture claims to tolerate should be sufficient, through the remaining trust, dependency, and authority structure, to derive a capability or state in H.
This is stronger than requiring individual components to be hardened. It asks whether the composition of the remaining architecture still permits high-consequence authority to emerge after a plausible foothold has already been obtained.
The architectural problem is consequently one of derivation interdiction: defensive controls should be placed where they make the inference structures leading to H incomplete. Because the model is based on prerequisite sets rather than simple pairwise graph edges, a high-consequence derivation may require several capabilities simultaneously. Breaking any indispensable prerequisite can therefore be sufficient to make that derivation infeasible.
Control selection as a constrained design problem
Suppose each candidate control d\in\mathcal{D} has an implementation cost c(d)\geq0. The term cost can represent more than acquisition price. Depending on the analysis, it can include engineering effort, lifecycle burden, operational complexity, outage requirements, maintenance effort, or another explicitly defined decision metric.
Let \mathfrak{K}_{\mathrm{adm}}\subseteq2^{\mathcal{D}} denote the family of control sets that are technically and operationally admissible. An idealized control-selection problem is then
This formulation is intentionally idealized. It is not a complete quantitative cyber-risk model, and the objective should not be interpreted as though security engineering reduced to summing acquisition costs. Real defensive controls have:
uncertain effectiveness;
implementation interactions;
shared dependencies;
operational side effects;
maintenance requirements;
failure modes;
lifecycle constraints;
safety and availability consequences.
Those effects determine both the feasible family \mathfrak{K}_{\mathrm{adm}} and, in a more detailed model, the definition of the cost function itself.
The value of the formulation is architectural rather than predictive. It makes explicit that the design objective is not to harden every asset equally. It is to identify a feasible combination of controls for which none of the modeled initial footholds can still derive a high-consequence state.
Control leverage depends on which derivations it breaks
The model also explains why controls that appear very different technologically can have similar architectural value.
Consider the Polish private-APN incident. Several controls could have interrupted the reconstructed derivation:
APN peer isolation could have prevented the compromised remote site from acquiring reachability to the CHP-connected WAGO controller;
replacement of default WAGO credentials could have prevented reachability from becoming administrative authority;
disabling or tightly constraining SSH could have prevented the WAGO controller from becoming a tunnel endpoint;
stronger segmentation between the industrial edge and the CHP control environment could have prevented the pivot from extending reachability into the OT domain;
stronger controller authorization could have prevented PLC reachability from becoming programming or operating-mode authority.
These mechanisms operate on different technologies and at different architectural locations. In the inference model, however, each acts by invalidating at least one prerequisite or architectural condition required to derive a later capability. For example, consider a rule r=\left(P_r,v_r,\tau_r\right) that is required by every derivation from a given initial foothold C_0 to the high-consequence set H. If a selected control set K makes q_r^K(z)=0, then that rule can no longer contribute to any such derivation.
The architectural leverage of a control therefore depends not simply on how many assets it protects, but on which high-consequence derivations depend on the capability or condition it constrains. A control protecting one strategically positioned trust relationship can consequently provide greater security benefit than uniform hardening of many peripheral devices.
Multiple controls can interrupt the same derivation structure
The Polish incident also illustrates that there is rarely one uniquely correct control location. Peer isolation, credential hardening, tunnel prevention, segmentation, and controller authorization constrain different prerequisites or feasibility conditions within the same derivation structure.
This is desirable. If a high-consequence derivation can be interrupted only by one control, that control becomes a critical security dependency. Defense in depth instead seeks several opportunities for interruption whose failure modes are sufficiently independent. The relevant architectural question is therefore not merely:
Where can this high-consequence derivation be interrupted?
It is:
Which independent controls make the required capability derivation fail, and does the architecture still prevent high-consequence closure if one of those controls is lost?
The same closure model can represent this stronger requirement. For control set K and one failed control d \in K, the surviving defensive set is K\setminus\{d\}. A single-control-failure tolerance requirement can therefore be written as C_{z,K\setminus\{d\}}^{*}(C_0)\cap H=\varnothing for every control failure, foothold, and operational state that the design explicitly claims to tolerate. More generally, let \mathfrak{L}\subseteq 2^K denote the family of defensive-loss scenarios against which the architecture is intended to remain resilient. Each L\in\mathfrak{L} is a set of controls assumed unavailable, ineffective, or compromised. The robust high-consequence separation requirement is then
C_{z,K\setminus L}^{*}(C_0)
\cap
H
=
\varnothing
\qquad
\forall C_0 \in \mathfrak{C}_0,
\;
\forall L \in \mathfrak{L}.
\tag{10}
This gives defense in depth a stronger formal meaning. It is not simply the presence of multiple security mechanisms.
It is the property that the loss of one or more explicitly modeled defensive mechanisms does not restore a complete capability derivation from a plausible foothold to an unacceptable industrial consequence. The closure model therefore turns control placement into a concrete architectural question:
Which controls must remain effective so that every modeled derivation from a plausible initial compromise to a high-consequence state contains at least one infeasible inference step?
Consequence concentration
The capability-inference model also explains why vulnerability severity alone is an incomplete prioritization mechanism in OT.
A technically vulnerable device may have limited operational significance if the capabilities derivable from its compromise remain read-only or tightly contained. Conversely, an engineering workstation, remote-access gateway, identity service, carrier relationship, or shared management platform may deserve urgent remediation even when no single vulnerability affecting it appears exceptional, because many independent derivations of high-consequence authority depend on it.
The relevant structural property is therefore not simply vulnerability severity, but consequence concentration: the extent to which different derivations of unacceptable authority or consequence depend on the same architectural element.
Minimal high-consequence derivations
The capability model developed above is based on inference rules with potentially multiple prerequisites. A conventional graph path is therefore too restrictive as the fundamental object of analysis.
Let \mathcal{R}'\subseteq\mathcal{R} be a subset of inference rules. Let C_{z,\mathcal{R}'}^{*}(C_0) denote the capability closure obtained from initial foothold C_0 when only rules in \mathcal{R}' are available. For C_0 \in \mathfrak{C}_0 and h \in H, a rule set \mathcal{R}' is a minimal derivation of h from C_0 when h\in C_{z,\mathcal{R}'}^{*}(C_0), but h\notin C_{z,\mathcal{R}''}^{*}(C_0)\qquad\forall\mathcal{R}''\subsetneq\mathcal{R}'. In other words, the rules in \mathcal{R}' are jointly sufficient to derive the high-consequence state, and removing any one or more indispensable rules from that particular derivation makes it incomplete.
Let \mathfrak{D}_{\min}\left(\mathfrak{C}_0,H;z\right) denote the set of all such minimal derivations across the modeled initial-compromise scenarios and high-consequence states. A member \delta\in\mathfrak{D}_{\min}\left(\mathfrak{C}_0,H;z\right) can therefore be represented as \delta=\left(C_0,h,\mathcal{R}_{\delta}\right), where \mathcal{R}_{\delta} is the minimal rule set required to derive h from C_0 under state z.
This construction is more general than an ordinary attack path. A derivation may branch and may require several capabilities simultaneously before a later capability becomes available.
Architectural elements can concentrate consequence
The rules themselves depend on concrete architectural elements.
Let \mathcal{X} denote the set of security-relevant architectural elements, including, where relevant:
endpoints;
engineering workstations;
gateways;
controllers;
identity services;
remote-access systems;
carrier relationships;
trust boundaries;
shared management platforms;
protocol services;
supplier connections.
For each inference rule r, let \operatorname{dep}(r)\subseteq\mathcal{X} denote the architectural elements on whose state or configuration the feasibility of that rule depends. For example, a rule that derives cross-site reachability may depend on a particular APN and its peer-isolation policy. A rule that derives controller administration may depend on a reachable controller service and its authentication configuration. A rule that derives programming authority may depend on an engineering workstation, controller mode, and authorization mechanism.
For architectural element x \in \mathcal{X}, define its high-consequence derivation participation as
B_H(x;z)
=
\left|
\left\{
\delta
\in
\mathfrak{D}_{\min}
\left(
\mathfrak{C}_0,
H;
z
\right)
\;\middle|\;
\exists r
\in
\mathcal{R}_{\delta}
:
x
\in
\operatorname{dep}(r)
\right\}
\right|.
\tag{11}
B_H(x;z) counts the modeled minimal high-consequence derivations whose feasibility depends on architectural element x. It is not a probability, an expected loss, or a complete risk score. It is a structural measure of consequence concentration.
An engineering workstation with high B_H(x;z) represents a concentration of trusted authority because many otherwise distinct derivations depend on it. A private APN with high B_H(x;z) represents a common propagation mechanism when several high-consequence derivations require the reachability it provides.
An identity service with high B_H(x;z) may constitute a common root of authority across several operational environments. Likewise, a control capable of invalidating rules occurring in many minimal derivations can have high architectural leverage even if it directly protects relatively few devices.
The measure is also representation-dependent. Splitting one architectural condition into several inference rules, or modeling the same system at a different level of abstraction, can change the numerical count. B_H(x;z) should therefore be used to compare structural concentration within a consistently constructed model, not as an absolute cross-organization risk metric.
Its purpose is prioritization: to identify where compromise, misconfiguration, or loss of one architectural element participates repeatedly in the derivation of unacceptable consequence. This is why simple counts of vulnerabilities, exposed IP addresses, or affected PLCs can misrepresent actual OT risk.
How the taxonomy relates to the capability-inference model
The capability-inference model formalizes one important part of the taxonomy: how an initial foothold can acquire additional capabilities and eventually make an unacceptable operational state derivable. It should not be interpreted as a complete mathematical representation of every OT risk mechanism.
The sixteen failure classes affect the cyber-physical system in different ways.
Failure class
Effect on the capability-inference system
Asset knowledge
Makes the defender’s representation of capabilities, rules, architectural dependencies, or feasibility conditions incomplete
Organizational ownership
Leaves end-to-end high-consequence derivations without accountable governance; it does not necessarily change the technical closure directly
Exposure
Makes unnecessary reachability rules feasible
Trust boundary
Allows capabilities to be derived across domains that were expected to constrain propagation
Shared infrastructure
Creates common architectural dependencies on which multiple derivations rely
Identity
Makes authentication-related derivations possible without sufficiently attributable or trustworthy identity
Privilege
Makes capabilities derivable beyond those required by the principal’s legitimate operational role
Third-party authority
Introduces externally controlled prerequisites or rules whose scope, duration, or downstream authority exceeds the maintenance mission
Protocol semantics
Allows an authorized communications relationship to derive industrial operations whose authority is broader than the network policy expresses
Engineering workstation
Concentrates programming tools, credentials, trusted applications, and controller relationships that can satisfy prerequisites for high-authority derivations
Controller integrity
Prevents reliable determination that the controller state appearing in the model corresponds to an independently approved state
Process observability
Weakens the independent evidence available to determine whether digital state corresponds to physical state
Monitoring
Reduces the ability to observe derivation of attacker capabilities and to trigger defensive changes to architectural state z
Safety boundary
Allows ordinary-control compromise to satisfy prerequisites affecting mechanisms intended to bound physical consequence
Availability dependency
Provides derivations to operational consequence that do not require direct acquisition of controller authority
Recovery integrity
Allows attacker capability or untrusted state to remain present in the post-recovery architecture
Table 4: Relationship between the sixteen failure classes and the capability-inference model. Not every failure class enlarges attacker closure directly: some affect model completeness, observability, consequence containment, or reconstitution of trust.
Several distinctions follow:
Reachability, identity, privilege, protocol, third-party-authority, and trust-boundary failures primarily change which capability-inference rules are feasible or which prerequisites can be satisfied.
Shared-infrastructure failures create common dependencies across otherwise distinct derivations, increasing the possibility of correlated or common-mode compromise.
Controller-integrity and process-observability failures weaken the defender’s ability to establish that the digital and physical states represented in the analysis are trustworthy.
Monitoring failures do not necessarily enlarge the static closure directly. Instead, they reduce the defender’s ability to observe capability acquisition and to respond by changing z before a high-consequence state becomes derivable.
Safety-boundary and availability-dependency failures determine which routes from digital compromise to operational consequence exist and how much physical consequence remains possible after ordinary controls fail.
Recovery-integrity failures concern the state produced after containment and restoration: a nominally recovered system may still contain attacker capability or untrusted cyber-physical state.
Asset-knowledge and organizational-ownership failures operate at the level of the defender’s model and governance. The actual architecture may contain a dangerous derivation even though the defender has failed to represent it or assign responsibility for it.
The taxonomy is therefore broader than the capability-inference model. The formal model represents the derivation of attacker capability and consequence. The taxonomy describes the larger cyber-physical system within which those derivations are created, hidden, constrained, amplified, governed, and eventually recovered from.
The practical assessment problem
The combined taxonomy and capability model reduce an OT architecture review to four principal questions.
Which plausible initial compromises can derive unacceptable operational capabilities or process states?
Which prerequisites, architectural conditions, and inference rules are required for those high-consequence derivations?
Which components, identities, shared services, trust relationships, and other architectural elements recur across multiple minimal derivations?
Which technically and operationally admissible controls invalidate the required rules or prerequisites so that no modeled high-consequence derivation remains feasible?
This produces more useful architectural statements than exposure counts alone. Instead of:
37 controllers are reachable.
the analysis should be able to establish:
Under the current trust relationships, compromise of the vendor-access gateway is sufficient to satisfy the prerequisites required to derive programming authority over four PLCs.
Instead of:
The sites use a private APN.
it should determine:
Compromise of any one remote site creates network reachability to three otherwise independent facilities because peer isolation is absent.
Instead of:
The plant has an engineering workstation.
it should establish:
This engineering workstation is an architectural dependency of every currently identified minimal derivation of controller-programming authority from the modeled external footholds.
And instead of:
The PLC network is segmented from the safety system.
it should determine whether:
Any modeled derivation permits capabilities acquired in ordinary control to satisfy the prerequisites required to modify safety-relevant functions.
That is the practical purpose of combining the taxonomy with the capability-inference model.
The primary security object is not the number of vulnerable assets, nor even the number of reachable assets. It is the structure of minimal capability derivations through which plausible digital compromise can acquire progressively stronger operational authority and ultimately make unacceptable cyber-physical states reachable.
A system may contain hundreds of reachable telemetry endpoints yet expose few derivations to high-consequence control. Another system may contain only one remotely reachable engineering environment while making that environment an indispensable dependency in several short derivations of programming authority across an entire process area. The stronger OT security question is therefore not merely:
What can an attacker reach?
It is:
What operational authority or physical consequence can be derived from a plausible foothold, which architectural prerequisites make that derivation possible, where is consequence concentrated across otherwise distinct derivations, and which independent controls make every modeled derivation of an unacceptable cyber-physical state infeasible?
That shift, from enumerating vulnerable assets to reasoning about the derivation, concentration, and interruption of operational authority, is the principal analytical contribution of the taxonomy and the capability-inference model.
Why OT security is not enterprise IT security
The statement that OT is different is often reduced to a reversal of the confidentiality-integrity-availability triad. That is not a sufficient model.
NIST notes that safety, availability, integrity, and confidentiality can assume different priorities in OT than in conventional enterprise IT, while their relative importance remains dependent on the process and mission.31 The international Principles of Operational Technology Cyber Security make the more fundamental point: safety is paramount, because cyber failure can propagate into harm to people, equipment, the environment, or essential services.32
The deeper distinction is therefore not a reordered security triad. It is that OT participates directly in a physical process. In enterprise computing, failure of a cybersecurity mechanism generally affects information, access, or a digital service. In OT, the same mechanism can also alter sensing, control, timing, actuation, process availability, or the operator’s ability to place the plant in a safe state. A cybersecurity mechanism can therefore become part of the cyber-physical system whose safety and availability it is intended to protect.
Security controls become part of the physical failure model
Consider a controlled physical process whose nominal dynamics are
\tau_0 is the delay already present in the sensing, computation, communication, and actuation path.
Now introduce cybersecurity control c: for example, a firewall, protocol-inspection gateway, authentication proxy, cryptographic gateway, endpoint agent, or access broker.
One possible effect of the control is additional timing variation. Inspection, queueing, retransmission, authentication, policy evaluation, degraded operation, and failover can make the additional delay time-dependent.
Let \tau_c(t)=\tau_0+\delta_c(t), where \delta_c(t) represents the timing perturbation introduced by control c. The corresponding process model becomes
This model isolates timing as one mechanism through which a cybersecurity control can affect the physical process. A real implementation may additionally introduce packet loss, failover states, resource exhaustion, service unavailability, changed routing, or other failure behavior. Those effects must be included in the engineering model when they are relevant to the process.
Let \mathcal{S} denote the acceptable safe region of the physical state space. Let \mathfrak{D} denote the disturbance trajectories included in the design basis, and let \Delta_c denote the family of timing perturbations that control c is assumed to produce in normal, degraded, and failover operation. A basic safety requirement is
Whenever deployment or failure of a cybersecurity mechanism can change whether this property holds, that mechanism belongs inside the process-safety and reliability argument rather than outside it. This is the first fundamental difference from ordinary enterprise security engineering:
A cybersecurity control in OT can simultaneously constrain attacker capability and modify the physical system’s own failure behavior.
The two effects must be analyzed independently.
Cyber benefit is a property of the resulting capability closure
The capability-inference model developed earlier provides a precise way to describe the cyber side of the problem. Let C_z^{*}(C_0) denote the capability closure generated from initial foothold C_0 under architectural state z, and let H\subseteq V be the high-consequence capability and state set. Define the high-consequence portion of the closure as
\mathcal{H}_z(C_0)
=
C_z^{*}(C_0)
\cap
H.
\tag{15}
Suppose deployment of control c changes the architecture from state z to state z_c. Let the design-basis foothold set be \mathfrak{C}_0=\{C_0^{(1)},\ldots,C_0^{(m)}\}. A strong definition of unambiguous cyber improvement requires that the set of high-consequence capabilities derivable after deployment of the control is never larger than before, for any modeled foothold:
To constitute a strict improvement rather than merely a non-degradation, the control must additionally reduce that set for at least one design-basis foothold:
Together, Equation 16 and Equation 17 define a control as unambiguously beneficial in the modeled cyber architecture: it introduces no additional high-consequence capability for any design-basis foothold and removes at least one such capability for at least one foothold.
This formulation matters because a security product does not necessarily only remove attacker capabilities. A remote-access broker may eliminate broad routed access while introducing a new identity or management dependency. A centralized security gateway may block dangerous protocol operations while becoming a common administrative target. A ZTNA service may reduce lateral movement while creating dependence on a policy engine or external service. If a control removes one high-consequence derivation but introduces another, the post-control high-consequence set need not be a subset of the original one.
The model therefore evaluates the resulting architecture, not the nominal strength of the security product. The strongest design objective remains high-consequence separation, meaning that no design-basis foothold can derive any capability or state contained in the high-consequence set:
Equation Equation 18 is stronger than the improvement conditions above: a control may reduce the set of derivable high-consequence capabilities without eliminating it completely, whereas high-consequence separation requires that the post-control architecture make every modeled high-consequence capability underivable from every design-basis foothold.
But cyber improvement alone is not sufficient to make a control acceptable in OT. The same mechanism must also preserve the process-safety and mission properties required by the plant:
A deep-packet-inspection gateway can remove dangerous protocol operations while becoming a common dependency for control traffic.
An authentication broker can remove shared credentials while making privileged maintenance dependent on an identity service.
Endpoint protection can prevent malicious tooling while interfering with an engineering application.
A cryptographic gateway can strengthen peer identity while introducing key-management, timing, and failover dependencies.
A technically stronger cybersecurity mechanism is therefore not automatically a stronger OT design.
If a control removes one attacker derivation but introduces a failure mode whose loss eliminates an essential control function, the architecture may have exchanged one high-consequence risk mechanism for another. The design objective is consequently not to maximize the number or nominal strength of security controls. It is to reduce derivable high-consequence authority while preserving the process invariants, trusted control capability, and failure behavior required by the physical mission.
This can justify different mechanisms at different architectural positions. Strong identity enforcement, protocol inspection, session mediation, and contextual access control may be appropriate at remote-access boundaries, engineering conduits, and supervisory interfaces where additional processing and dependency can be tolerated. A tightly coupled control loop may instead require deterministic network enforcement, controller-native protection, hardware-supported mechanisms, or security controls positioned outside the time-critical feedback path.
The relevant criterion is not whether a security mechanism is intrinsically strong. It is whether, in its actual architectural position, it removes unacceptable attacker capability without creating another unacceptable route to loss of physical safety or mission capability.
Availability means availability of the physical mission
The same reasoning changes the meaning of availability:
Server uptime is not industrial availability.
A supervisory server can be unavailable while autonomous controller logic and local operators retain sufficient trusted capability to maintain the process safely.
Conversely, every digital service can remain responsive while a compromised controller drives the physical process toward an unacceptable state.
Availability should therefore be defined relative to the physical mission and the trusted control capability that survives under the current cyber-operational state.
Let z denote the cyber-operational state of the plant: which controllers, networks, supervisory functions, identities, remote services, engineering systems, communications paths, and local-control mechanisms remain available and trustworthy. Let \mathcal{U}(z) denote the set of instantaneous control actions that remain technically available through trusted operational paths under state z.
Because safe operation depends on sequences of actions rather than isolated commands, let \mathfrak{U}(z) denote the corresponding set of admissible control trajectories that can actually be realized through those trusted paths.
The distinction is important: a controller that remains online but whose running state cannot be trusted should not automatically contribute to \mathcal{U}(z) merely because it responds to network requests. Likewise, a remote engineering service that is technically reachable but whose identity path has been compromised should not automatically be treated as trusted operational capability.
Let \mathfrak{M} denote the set of physical state and control trajectories that satisfy the required industrial mission over a defined mission horizon [0,T_{\mathcal{M}}]. For cyber-operational state z, define the mission-viability set as
The set \mathcal{K}_{\mathcal{M}}(z) contains the initial physical states from which the required mission can still be maintained safely using only control trajectories that remain trustworthy and realizable under cyber state z, for the disturbances included in the design basis.
This is a stronger and more operationally useful definition of availability than the uptime of any individual server or network service. If the cyber-operational state varies with time, mission availability can be represented as
where \mathbf{1}[\cdot] is the indicator function.
This separates component availability from mission availability. Consider a supervisory outage that changes the cyber state from z to z'. The surviving trusted action set may contract \mathcal{U}(z')\subseteq\mathcal{U}(z). Yet if x(t)\in\mathcal{K}_{\mathcal{M}}(z'), the plant has lost digital capability without losing the physical mission. Local autonomous control, manual operation, or another trusted degraded mode has preserved mission viability.
The converse is equally important. Suppose every major server remains online but controller integrity can no longer be established. If x(t)\notin\mathcal{K}_{\mathcal{M}}(z), then the required mission is no longer demonstrably viable under the current trusted-control state, even though conventional infrastructure monitoring may report high availability.
Resilience should therefore be evaluated in terms of mission-essential trusted capability, not merely digital uptime. Local autonomous control, manual fallback, independent protection, redundant communications, islanded operation, and graceful degradation are valuable because they can preserve or enlarge \mathcal{K}_{\mathcal{M}}(z) when cyber failure removes part of the normal control architecture.
For each credible degraded cyber state z', the architecture should therefore determine:
Which control actions remain trustworthy and available in \mathcal{U}(z')?
Which control trajectories remain realizable in \mathfrak{U}(z')?
Which physical states remain inside \mathcal{K}_{\mathcal{M}}(z')?
Can the required mission continue safely?
If full mission continuation is impossible, can the process reach an approved safe degraded or shutdown state before leaving \mathcal{S}?
The relevant availability question is therefore not:
How many digital services remain online?
It is:
Are the surviving trusted control capabilities sufficient to keep the physical mission viable?
A security control must satisfy both cyber and mission constraints
The two models can now be combined without treating cyber risk and process risk as additive quantities. For a candidate control c, the cyber model evaluates the change from \mathcal{H}_z(C_0) to \mathcal{H}_{z_c}(C_0). The process model evaluates whether the behavior introduced by c preserves the safe set \mathcal{S}. The mission model evaluates whether the resulting cyber-operational state leaves the relevant physical operating states inside \mathcal{K}_{\mathcal{M}}(z_c).
A control is therefore OT-admissible only if three logically distinct claims can be supported:
Cyber improvement: it removes or prevents high-consequence capability derivations without introducing new unacceptable ones.
Physical admissibility: its normal, degraded, and failure behavior preserves the required process-safety properties.
Mission admissibility: the resulting architecture retains sufficient trusted control capability to perform the required mission or reach an approved degraded state.
These are not interchangeable properties: a control can satisfy the first and fail the second, or satisfy the second and fail the third. Conversely, a highly available control that provides no meaningful reduction in derivable attacker authority may add complexity without providing the intended cyber benefit.
This three-part test is the deeper reason OT cybersecurity cannot be implemented by transplanting enterprise controls into industrial systems without process-specific engineering analysis.
Fail closed is not a universal requirement
The same distinction applies to failure behavior. In many information systems, denying access when a security mechanism fails is a reasonable default.
In OT, the correct response depends on the physical process. Closing one valve may terminate the release of hazardous material. Closing another may create overpressure. Stopping one pump may prevent mechanical damage. Stopping another may remove essential cooling. The correct requirement is therefore not generically fail open or fail closed.
Let t_f denote the time at which the security mechanism fails. Let \mathcal{S}_f\subseteq\mathcal{S} denote the approved safe or degraded target region associated with that failure scenario, and let T_f be the maximum acceptable transition time. An acceptable failure response must preserve safety during the transition:
x(t)
\in
\mathcal{S}
\qquad
\forall t
\in
[t_f,t_f+T_f].
\tag{21}
Where the process design requires transition to a defined safe or degraded state, it should additionally satisfy
x(t_f+T_f)
\in
\mathcal{S}_f.
\tag{22}
Whether the correct response is to open, close, hold, isolate, transfer to local control, continue autonomously, or trip the process is therefore determined by process engineering and functional safety.
Cybersecurity defines the failure scenario. Process engineering determines the physically admissible response.
Patching introduces controlled change risk
Vulnerability remediation exposes another important difference between enterprise and industrial systems.
Leaving a known vulnerability unresolved can preserve an attacker inference rule that should be made infeasible. Applying an inadequately tested patch or firmware update can instead introduce:
incompatibility;
changed timing;
controller restart;
communication loss;
altered process behavior;
changed protocol semantics;
invalidated supplier assumptions;
invalidated safety assumptions.
These are different failure mechanisms and should not be collapsed into a pseudo-quantitative equation such as vulnerability risk plus change risk. They need not be independent, commensurable, or additive. Patching should instead be evaluated using the same architectural criteria established above.
A patch is valuable when the resulting state z_{\mathrm{patch}} removes unacceptable attacker capability, for example by invalidating an inference rule associated with a known vulnerability. But deployment is operationally acceptable only if the resulting system also preserves:
the safety invariance required for the physical process;
the trusted control trajectories required by the mission;
the relevant mission-viability region;
the ability to recover if the change fails.
Testing, supplier validation, configuration backup, staged rollout, representative test environments, maintenance windows, and rollback procedures reduce uncertainty associated with the change. They do not justify indefinite acceptance of a known vulnerability. The appropriate engineering question is therefore not:
Is patching safer than not patching?
It is:
Which remediation strategy removes the unacceptable capability derivation while preserving the safety, mission, and recovery properties required by the plant?
For some assets, that strategy will be immediate patching. For others, it may require a maintenance outage, staged firmware migration, replacement of adjacent equipment, temporary compensating controls, or migration of an entire process cell.
The problem is amplified by asset longevity. NIST notes that OT systems commonly remain in service substantially longer than conventional enterprise computing platforms.33 A sustainable OT security architecture must therefore support vulnerability remediation and controlled technical change throughout an industrial lifecycle, not merely at commissioning.
Endpoint controls cannot always be installed on endpoints
Many PLCs, RTUs, drives, protection devices, sensors, and other embedded industrial systems cannot safely support conventional endpoint agents, arbitrary security software, or intrusive vulnerability scanning. Their operating systems may be proprietary or constrained, their computing resources may be tightly dimensioned, vendor support may prohibit unvalidated software, and active testing may interfere with communications or deterministic behavior.
That limitation does not remove the underlying security objective; it changes where and how the required security property must be enforced. Depending on the asset and threat scenario, protection may instead be distributed across controller-native authentication and authorization, service and protocol restriction, dedicated engineering workstations, constrained network conduits, passive monitoring, protocol-aware gateways, application allowlisting on supported hosts, physical-access controls, and procedural mechanisms governing maintenance or programming.
The relevant architectural principle is therefore broader than endpoint protection. If an endpoint cannot itself enforce a security property required by the risk model, the architecture must either provide that property through another trustworthy mechanism or explicitly accept the residual capability that remains derivable without it. The requirement does not disappear merely because the target technology cannot implement the control locally.
This follows directly from the capability-inference model developed earlier. A defensive mechanism is valuable when it makes an unacceptable capability derivation infeasible; the mechanism need not reside on the asset at which the final capability would otherwise be exercised. A conduit can remove unnecessary reachability, a privileged-access architecture can prevent administrative authority from being derived, an engineering workstation can constrain programming capability, and an independent monitoring system can make misuse observable even when the controller itself provides little native security functionality. The design question is consequently not:
Can an endpoint-security product be installed on this device?
It is:
Which security properties must hold around this device, which of them can the device enforce natively, and where must the remaining properties be implemented so that the high-consequence derivations identified by the architecture remain infeasible?
Containment can alter the physical process
Incident containment also differs fundamentally from ordinary enterprise practice because isolation of an OT component can remove not only attacker connectivity but also legitimate sensing, control, engineering, supervision, or protection capability. Disconnecting a compromised office endpoint usually reduces digital attack capability while leaving the underlying business process conceptually intact; disconnecting an HMI, PLC, engineering workstation, safety gateway, telecommunications link, or remote station may change what the physical process can observe, what commands can be issued, which autonomous functions remain available, and whether operators can still drive the plant toward a safe state.
Containment is therefore itself a cyber-physical intervention. Let z denote the current cyber-operational state and let an incident-response action a produce a new state z_a.
The action can change the trusted control capabilities represented by \mathcal{U}(z_a) and the corresponding set of realizable trusted control trajectories \mathfrak{U}(z_a). There is no general requirement that these sets be strict subsets of their pre-containment equivalents. Isolation commonly removes capabilities, but an incident-response action may also activate local control, manual fallback, an independent communications path, or another degraded operating mode. The relevant quantity is therefore not how many digital capabilities disappear, but whether the resulting cyber-operational state remains compatible with the physical mission.
For a containment action intended to preserve the mission, the current physical state must remain viable under the resulting architecture x(t_a)\in\mathcal{K}_{\mathcal{M}}(z_a), where t_a is the time at which the containment action takes effect and \mathcal{K}_{\mathcal{M}}(z_a) is the mission-viability set defined previously.
If this condition does not hold, containment may still be correct, but the objective must change from mission continuation to controlled transition toward an approved degraded or shutdown state. In that case, the incident-response decision must be coordinated with the process response rather than evaluated as a purely digital isolation action.
Before isolating an OT component, incident response must therefore establish which process functions depend on it, which trusted control actions will disappear, which autonomous functions remain active, whether local operators retain sufficient authority, whether independent protection remains available, and what physical transition the plant is expected to undergo after the isolation occurs. These are not secondary operational considerations; they determine whether the proposed containment action is itself safe. The correct question is consequently not merely:
How quickly can this system be disconnected?
It is:
What trusted control capability will remain after disconnection, and what will the physical process do under that resulting state?
This distinction also explains why pre-engineered degraded modes are important in OT incident response. Local control, autonomous fallback, tested isolation procedures, independent communications, manual authority, and known shutdown sequences reduce the need to improvise the physical consequence of a cybersecurity action during an active incident.
Safety independence must survive cyber compromise
Redundancy is not equivalent to independence.
Two controllers can be redundant against random hardware failure while remaining dependent on the same engineering workstation, credentials, identity infrastructure, switching or routing equipment, remote-maintenance path, firmware repository, or privileged administrators. Such an architecture may tolerate failure of one controller while remaining vulnerable to a single cyber compromise that acquires authority over both.
The relevant security property is therefore not the number of redundant devices but the degree to which ordinary control and independent protection remain separated under plausible cyber compromise.
Let D_C denote the set of cyber dependencies whose compromise can materially influence the ordinary control system, and let D_S denote the corresponding dependency set for the safety or independent protective system. Their shared dependency set is
D_{\mathrm{shared}}
=
D_C
\cap
D_S.
\tag{23}
The intersection identifies candidate common-mode cyber dependencies. Membership in D_{\mathrm{shared}} does not by itself prove that compromise of that dependency will defeat both systems; the actual consequence still depends on the authority provided by the dependency, the controls surrounding it, and the failure behavior of the two systems. It does identify precisely where an independence claim requires additional justification.
Complete disjointness, D_C\cap D_S=\varnothing, would provide a particularly strong form of architectural separation, but it is not a universal requirement and may be impractical in real plants. Shared networking infrastructure, maintenance personnel, time sources, engineering processes, or physical facilities can sometimes be justified. The engineering objective is instead to identify the shared dependencies explicitly, minimize unnecessary sharing, and ensure that no shared high-authority dependency silently collapses the independence on which the protection concept relies.
A shared read-only monitoring path is therefore not equivalent to a shared engineering workstation capable of modifying both systems. A common physical cabinet is not equivalent to a common privileged identity. The importance of D_{\mathrm{shared}} depends on what authority compromise of each shared dependency can derive.
This connects safety independence directly with the capability-inference model. If compromise of some shared dependency d\in D_{\mathrm{shared}} can satisfy prerequisites for both ordinary-control authority and modification of the mechanism intended to constrain that authority, then the safety architecture contains a cyber common-mode path that must be explicitly addressed. The relevant question is consequently not merely:
Are the control and safety systems redundant or nominally separated?
It is:
Which cyber dependencies can influence both systems, what authority does each shared dependency provide, and can compromise of one of them defeat both ordinary control and the mechanism intended to bound its physical consequence?
Table 5 summarizes the engineering distinctions developed in this section.
Property
Enterprise IT tendency
OT engineering implication
Primary consequence
Information loss, fraud, privacy breach, or digital-service disruption
Cyber failure can propagate into injury, equipment damage, environmental harm, or loss of essential physical service
Timing
Predominantly soft real-time requirements
Some sensing, control, and protection functions depend on bounded latency, jitter, sequencing, or deterministic behavior
Availability
Availability of applications and digital services
Viability of the physical mission using trustworthy surviving control capability
Security-control failure
Denial of access is often an acceptable default
Failure behavior must be derived from process-safe and mission-viable behavior
Patching
Frequent standardized deployment is often feasible
Remediation requires process-aware validation, scheduling, rollback, and lifecycle management
Asset lifetime
Relatively short technology cycles
Industrial systems may remain operational for decades
Endpoint protection
General-purpose security agents are commonly feasible
Many embedded endpoints cannot safely support arbitrary agents or intrusive security software
Security testing
Active scanning is commonly tolerated
Some OT assets require passive discovery or tightly controlled active testing
Containment
Disconnecting a compromised endpoint usually reduces digital consequence
Isolation can itself remove trusted control capability and alter physical-process behavior
Recovery
Restore infrastructure, applications, identities, and data
Re-establish digital trust and reconcile the restored automation with the physical process state
Redundancy
Primarily supports digital-service continuity
Shared cyber dependencies can defeat nominally redundant control or protection systems simultaneously
Safety
Usually external to ordinary application-security analysis
Cyber controls, cyber failures, and shared dependencies can alter the process-safety argument itself
Table 5: Engineering differences that constrain how ordinary cybersecurity principles are applied to OT.
The conclusion is not that ordinary cybersecurity principles cease to apply. Authentication, least privilege, segmentation, vulnerability management, secure configuration, logging, monitoring, backup, recovery, secure development, and controlled third-party access remain fundamental, but their implementations must be evaluated against the physical mission rather than against digital risk alone. The stronger OT criterion is therefore:
Does the control reduce unacceptable derivable cyber authority while preserving the timing, mission viability, safety independence, failure behavior, and recoverability required by the physical process?
That engineering distinction provides the basis for the defensive architecture developed in the remainder of the article.
IEC 62443 as an architecture for industrial cyber risk
IEC 62443 is best understood as a lifecycle framework for assigning cybersecurity responsibilities, requirements, and capabilities across an industrial automation and control system, rather than as a single checklist, product certificate, or prescribed network topology.34
That distinction follows naturally from the preceding analysis. The failure taxonomy identified weaknesses in architectural knowledge, reachability, trust, identity, authority, control-path integrity, observability, safety, availability dependencies, and recovery. No single technical control can address all of those properties, because they arise at different lifecycle stages and are controlled by different actors. IEC 62443 consequently distributes responsibility across asset owners, system integrators and service providers, and product suppliers, while connecting organizational governance, system architecture, component capability, secure development, maintenance, and operational protection.
As of September 2026, the principal published parts relevant to this analysis include:
IEC 62443-2-1:2024: security program requirements for IACS asset owners;
IEC PAS 62443-2-2:2025: IACS Security Protection Scheme;
IEC TR 62443-2-3:2015: patch management in the IACS environment;
IEC 62443-2-4:2023: security program requirements for IACS service providers;
IEC 62443-3-2:2020: security risk assessment for system design;
IEC 62443-3-3:2013: system security requirements and security levels;
IEC 62443-4-2:2019: technical security requirements for IACS components.3536373839404142
The important architectural property of the series is composition. IEC 62443-2-1 places continuing security-program responsibility on the asset owner; IEC 62443-3-2 transforms system risk into architectural requirements through definition of the System under Consideration, zones, conduits, risk assessment, and target security levels; IEC 62443-3-3 specifies system security requirements and security levels; IEC 62443-2-4 addresses integration and maintenance processes performed by service providers; IEC 62443-4-1 governs secure product development; IEC 62443-4-2 defines technical component security requirements; IEC TR 62443-2-3 addresses patch management; and IEC PAS 62443-2-2 addresses the technical, physical, and procedural measures that compose the operational Security Protection Scheme.
No individual part is sufficient by itself. A secure component does not create a secure system merely by being installed, a segmented network does not compensate automatically for uncontrolled engineering authority, and a mature asset-owner program cannot make a legacy controller provide a native capability it does not possess. Missing properties must instead be addressed through system architecture, compensating measures, operational controls, replacement, or explicit residual-risk treatment.
Figure 4 summarizes the lifecycle relationship used in this article.
%%{init: {"theme": "neo", "look": "handDrawn", "layout": "elk"}}%%
flowchart TB
AO[Asset-owner security program<br/>IEC 62443-2-1]
RA[System risk assessment<br/>zones, conduits and SL-T<br/>IEC 62443-3-2]
SR[System security requirements<br/>IEC 62443-3-3]
SP[Integration and maintenance<br/>IEC 62443-2-4]
SDL[Secure product-development lifecycle<br/>IEC 62443-4-1]
CR[Component security capabilities<br/>IEC 62443-4-2]
PM[Patch-management process<br/>IEC TR 62443-2-3]
SPS[Operational Security Protection Scheme<br/>IEC PAS 62443-2-2]
AO -->|governance and risk ownership| RA
RA -->|risk-derived requirements| SR
SDL -->|secure development produces| CR
CR -->|component capability constrains design| SP
SR -->|system requirements guide| SP
SP -->|implemented architecture and procedures| SPS
AO -->|organizational and procedural controls| SPS
PM -->|lifecycle vulnerability treatment| SPS
SPS -->|validation, operation and reassessment| AO
Figure 4: Simplified relationship among the principal IEC 62443 responsibilities used in this article. Risk assessment establishes the System under Consideration, zones, conduits, and target security requirements; system requirements, component capabilities, secure product development, integration and maintenance, patch management, and asset-owner governance contribute to the operational Security Protection Scheme. Arrows represent lifecycle or requirement dependencies, not attacker-causal transitions.
This representation is intentionally simplified. IEC 62443 is not a linear waterfall process, and the arrows in the figure do not represent the attacker-capability semantics used in the incident diagrams. They express lifecycle and requirement dependencies: operational experience can change risk assumptions, system changes can alter zones or conduits, compensating measures can affect achieved capability, product updates can change the security properties available to integrators, and asset owners retain responsibility for reassessing the resulting system throughout its lifecycle.
The practical implication is that IEC 62443 should be applied as a system-of-systems assurance structure. Product capability, integration quality, architectural containment, operational procedure, maintenance, monitoring, and governance must compose into an effective security property at the deployed system level.
Zones and conduits turn risk into architecture
IEC 62443-3-2 requires the organization to define the System under Consideration (SuC), partition the relevant system into zones and conduits, assess risk, establish target security levels, and document the resulting security requirements.43
A zone groups assets that share security requirements, while a conduit groups communication channels connecting zones and subject to common security requirements. These abstractions are valuable because they force the architecture to express where trust changes, which communications are required, and where additional security enforcement is necessary.
For the architectural analysis developed here, represent the SuC as G_{\mathrm{SuC}}=\left(Z,\mathcal{C}\right), where Z=\left\{Z_1,Z_2,\ldots,Z_n\right\} is the set of security zones and \mathcal{C} is the set of conduits connecting them.
This is an architectural graph, not the capability-inference system introduced earlier. The distinction is fundamental. G_{\mathrm{SuC}} represents intended system decomposition and communication relationships; the capability-inference model represents what an attacker can derive when the actual properties of those zones, conduits, identities, services, and trust relationships are taken into account.
The two models nevertheless interact. A conduit that permits unnecessary communication can make reachability-related inference rules feasible; a conduit whose administration depends on a compromised identity system may fail to provide the assumed trust boundary; a conduit that permits a legitimate industrial protocol without constraining high-authority operations may correctly restrict network reachability while still leaving dangerous capability derivations possible.
To make the communication requirement precise, consider two zones Z_i and Z_j. Let \mathcal{F}_{ij}^{\mathrm{mission}} denote the set of communication flows required by the industrial mission, and let \mathcal{F}_{ij}^{\mathrm{allowed}} denote the flows permitted by the deployed architecture. A flow should be interpreted sufficiently specifically to distinguish at least source, destination, direction, and required service or protocol relationship. With that scope, a least-functionality design target is
A well-constrained conduit has \Delta\mathcal{F}_{ij}=\varnothing. This gives the zone-and-conduit model a direct connection to several classes in the earlier taxonomy. Exposure failure appears when unnecessary flows remain in \Delta\mathcal{F}_{ij}; trust-boundary failure appears when the conduit does not continue to constrain capability under the compromise scenarios it was intended to contain; shared-infrastructure failure appears when several zones depend on a common network, management plane, identity system, or other conduit dependency whose compromise can propagate authority across otherwise separate environments.
Network-flow minimization is nevertheless only one layer of the security problem. Suppose an engineering conduit legitimately permits an engineering workstation to communicate with an S7 controller. That relationship may be necessary for maintenance and therefore belong to \mathcal{F}_{ij}^{\mathrm{mission}}. Network authorization alone does not determine whether the same relationship may be used for diagnostic reads, project transfer, firmware operations, controller programming, or CPU mode changes.
The distinction corresponds directly to the earlier protocol-semantic failure. Zones and conduits constrain where communication can occur; they do not automatically constrain which industrial authority can be exercised through communication that is legitimately permitted.
IEC 62443 zoning should therefore be derived from risk, consequence, trust, security requirements, and operational dependency, rather than copied mechanically from an automation hierarchy such as the Purdue model. A safety system, privileged engineering environment, remote site, historian infrastructure, vendor-access environment, shared management service, or group of high-consequence controllers may justify a distinct zone even when those systems occupy the same nominal automation level.
The architectural test is not whether the network diagram contains enough boxes. It is whether the zoning and conduit structure makes the trust assumptions required by the risk assessment explicit and provides enforceable locations at which unnecessary capability derivations can be interrupted.
Security levels express required resistance, not architectural prestige
IEC 62443 security levels are often treated as though a larger number were simply a better cybersecurity score. That interpretation is misleading because a security level is not a generic maturity rating, a product-quality grade, or an architectural badge. It expresses the degree of resistance required or provided against progressively more capable adversaries, and it is tied to specific security requirements rather than to an abstract notion of overall security.44
The starting point is therefore risk, not the desire to assign the highest possible number. The required security capability should follow from the consequences being protected against, the threat assumptions established by the risk assessment, the role of the zone or conduit, and the operational constraints of the system. A high security level applied indiscriminately can be as poor an engineering decision as a low one if satisfying it introduces mechanisms that are incompatible with the timing, availability, maintainability, or safety properties of the physical process.
Three concepts must remain distinct:
SL-T (Target Security Level): the security capability required as an outcome of the risk assessment and system design process;
SL-C (Capability Security Level): the security capability that a system or component is capable of supporting;
SL-A (Achieved Security Level): the security capability actually realized in the deployed and configured system.
These terms describe different questions. SL-T asks what level of protection is required. SL-C asks what security capability is available from a system or component. SL-A concerns what the deployed architecture actually achieves after products, configuration, compensating measures, integration, and operational controls have been combined. A target therefore does not prove that the necessary capability exists; a capable component does not prove that the capability has been correctly deployed; and deployment of capable products does not by itself establish that the required system-level security properties have been achieved.
This distinction is particularly important because IEC 62443 is explicitly architectural. System-level security requirements need not be implemented independently by every component. Where a legacy controller lacks a native capability, the missing property may sometimes be provided through a compensating countermeasure elsewhere in the system, for example through a dedicated engineering environment, a constrained conduit, a protocol-aware gateway, physical-access restrictions, procedural controls, or other mechanisms permitted by the system design.45 The relevant question is whether the complete deployed system satisfies the applicable requirement, not whether every device implements the same mechanism locally.
The converse is equally important:
A component certificate or a high component capability level does not establish the achieved security level of the deployed system.
A controller may provide strong authentication while remaining reachable through an unnecessarily broad conduit. A gateway may support sophisticated access control while depending on a compromised identity plane. An engineering workstation may use hardened software while retaining authority over more controllers than its operational role requires. Security capability therefore becomes meaningful only through its composition with the rest of the architecture.
IEC 62443 organizes the relevant system security requirements around seven Foundational Requirements:
IAC (Identification and Authentication Control);
UC (Use Control);
SI (System Integrity);
DC (Data Confidentiality);
RDF (Restricted Data Flow);
TRE (Timely Response to Events);
RA (Resource Availability).
These dimensions should not be collapsed conceptually into one undifferentiated number. Different zones, conduits, and threat scenarios can require different emphasis across the seven Foundational Requirements. A remote-maintenance environment may depend particularly strongly on identification, authentication, use control, and restricted data flow; a safety-relevant control environment may place exceptional importance on system integrity and resource availability; a monitoring or evidence environment may depend heavily on integrity and timely response to events.
This is why the security-level concept is better understood as a risk-derived security profile than as a ranking of architectural prestige. The important engineering question is not whether a system can be described informally as “SL 3” or “SL 4”, but which requirements apply to each Foundational Requirement, why those requirements follow from the risk assessment, and how the deployed architecture demonstrates that they are satisfied.
The distinction also prevents a common misuse of SL-T, SL-C, and SL-A. They should not be treated as three numbers to be compared mechanically. IEC 62443 conformity is established through satisfaction of applicable requirements and requirement enhancements, supported by system architecture, component capabilities, configuration, compensating countermeasures, and evidence. A numerical label is a compact representation of that requirement structure; it is not a substitute for it.
For brownfield OT this point is especially important. A legacy device may prevent a desired property from being implemented natively, but that does not automatically make the target requirement irrelevant. The architecture must determine whether the property can be supplied elsewhere without creating an unacceptable new dependency, whether a compensating measure is sufficient, or whether the residual gap requires eventual replacement of the component. This is the same principle established earlier for endpoint controls: a missing local capability relocates the engineering problem; it does not eliminate it.
The security-level concept should therefore be used to express required resistance derived from risk, not to maximize nominal scores. A higher level is not automatically a better design if the mechanisms required to achieve it undermine the physical mission, while a high-capability certified component cannot compensate for weak zoning, excessive privileged access, uncontrolled maintenance paths, poor lifecycle governance, or loss of safety independence.
The connection with the capability-inference model developed earlier is then straightforward without requiring additional mathematical notation. IEC 62443 requirements provide architectural mechanisms that can make particular attacker capabilities no longer derivable: identification and authentication controls can prevent unauthenticated authority, use control can constrain privilege, restricted data flow can remove unnecessary reachability, system-integrity requirements can protect control logic and configuration, timely-response mechanisms can improve detection and intervention, and resource-availability requirements can preserve essential functions under attack.
Their value is therefore not demonstrated by the presence of a certificate or by the numerical level attached to an individual component. It is demonstrated when the resulting system architecture prevents the modeled initial compromises from accumulating into the high-consequence authority or states that the risk assessment identified as unacceptable.
The seven Foundational Requirements constrain different parts of the capability-inference system
IEC 62443-3-3 organizes its system requirements around seven Foundational Requirements (FRs).46 They should not be mapped one-to-one onto the sixteen failure classes developed earlier because the two classifications describe different objects. The Foundational Requirements organize normative system-security requirements, whereas the incident-derived taxonomy distinguishes architectural mechanisms through which cyber compromise can acquire operational significance.
The relationship is therefore many-to-many. Identity failure is addressed principally through Identification and Authentication Control but also depends on Use Control and System Integrity; protocol-semantic failure can involve Use Control, System Integrity, and Restricted Data Flow; safety-boundary failure cuts across several FRs because independence can depend simultaneously on identity, integrity, restricted communication, resource availability, and operational architecture. Asset-knowledge and organizational-ownership failures sit even further outside a direct FR mapping because they are addressed primarily through governance, risk assessment, architecture management, and lifecycle responsibility rather than through one technical requirement family.
Table 6 therefore identifies the principal relationships without claiming an exact equivalence between the two models.
Foundational Requirement
Principal architectural function
Principal taxonomy failures or concerns addressed
FR 1 — Identification and Authentication Control (IAC)
Establish and authenticate human, software, and device identities before authority is exercised
Identity failure; elements of third-party-authority failure
FR 2 — Use Control (UC)
Constrain which operations an authenticated principal may perform
Privilege failure; third-party-authority failure; elements of protocol-semantic failure
FR 3 — System Integrity (SI)
Preserve trustworthy software, firmware, configuration, communications, and system state
Controller-integrity failure; engineering-workstation failure; elements of protocol-semantic failure
FR 4 — Data Confidentiality (DC)
Protect information whose disclosure would be unauthorized or could assist subsequent compromise
Exposure of engineering, process, credential, and configuration information; no single corresponding failure class in the incident-derived taxonomy
FR 5 — Restricted Data Flow (RDF)
Restrict communication between zones, conduits, assets, and trust domains to required relationships
Exposure failure; trust-boundary failure; elements of shared-infrastructure failure
FR 6 — Timely Response to Events (TRE)
Make security-relevant events observable and support timely interpretation and response
Monitoring failure; elements of process-observability failure
FR 7 — Resource Availability (RA)
Preserve the resources and services required for continued or degraded operation under denial, exhaustion, or failure
Availability-dependency failure; aspects of resilience and recovery capability
Table 6: Principal relationships between IEC 62443-3-3 Foundational Requirements and the incident-derived failure taxonomy. The mapping is intentionally many-to-many because the two structures serve different analytical purposes.
The comparison also clarifies what each abstraction contributes. IEC 62443 provides normative requirements, capability expectations, lifecycle processes, and responsibilities; the taxonomy developed in this article emphasizes how weaknesses in those properties can compose into excessive reachability, derived authority, untrusted control state, impaired observability, loss of safety independence, or operational consequence. They are therefore complementary rather than competing models.
FR 5 is architecturally central, but insufficient alone
Restricted Data Flow is especially important because it determines which communication relationships can exist between zones and other security domains. In the capability-inference model developed earlier, FR 5 primarily influences whether prerequisites involving reachability can be satisfied; reducing unnecessary communication can therefore eliminate entire families of later derivations before identity, privilege, or controller authority become relevant.
Reachability control alone, however, is insufficient. FR 1 determines whether the architecture can establish who or what is acting; FR 2 constrains which capabilities an authenticated principal may exercise; FR 3 protects the integrity of the systems and states through which those capabilities acquire operational meaning; FR 4 limits disclosure of information that may assist later compromise or reveal sensitive process semantics; FR 6 determines whether abnormal acquisition or use of authority becomes observable and actionable; FR 7 preserves the resources needed to continue operation, degrade safely, or recover when attack or failure removes part of the normal system.
Suppose K denotes the collection of system requirements, component capabilities, compensating measures, integration controls, and operational mechanisms selected to realize the IEC 62443 design. The architectural objective remains the high-consequence separation condition introduced earlier:
C_{z,K}^{*}(C_0)
\cap
H
=
\varnothing
\qquad
\forall C_0\in\mathfrak{C}_0.
\tag{26}
This is not IEC 62443 notation. It is the capability-inference model developed in this article applied to an architecture whose controls have been derived using IEC 62443.
The equation gives the relationship a precise meaning: the resulting system should prevent every design-basis initial compromise from deriving a capability or state that belongs to the high-consequence set H. Different Foundational Requirements contribute to that result in different ways. Restricted Data Flow can eliminate unnecessary reachability; Identification and Authentication Control and Use Control can prevent reachability from becoming authenticated or privileged authority; System Integrity can prevent apparently legitimate control paths from operating on untrusted software or configuration; Timely Response to Events can permit intervention before capability accumulation reaches its terminal consequence; Resource Availability can preserve the mission even when part of the digital architecture has been attacked or removed.
This is also why the distinction between attack surface and derived operational authority remains important. A system can implement strong Restricted Data Flow while still exposing excessive authority through one deliberately permitted engineering relationship; conversely, a legacy controller with weak native security may still participate in a robust architecture if independent controls prevent compromise elsewhere from deriving high-consequence authority over it.
IEC 62443-3-3 itself does not substitute its functional requirements for a complete integrated security architecture.47 The actual architecture must be developed from the applicable requirements, risk assumptions, operational constraints, product capabilities, compensating measures, and lifecycle responsibilities of the specific system.
IEC 62443 therefore does not prescribe a particular ZTNA product, firewall vendor, universal DMZ topology, or mechanical reproduction of the Purdue model. It establishes security properties, responsibilities, risk-derived requirements, and capability expectations that the deployed architecture must collectively realize. The design question is consequently not:
Which IEC 62443 product should be installed?
It is:
Which zones, conduits, identities, privileges, component capabilities, compensating measures, monitoring functions, safety boundaries, and recovery mechanisms are required so that the target security properties survive realistic compromise without sacrificing the physical mission?
This interpretation is consistent with the central thesis developed throughout the article: industrial cybersecurity is not primarily the security of individual products. It is the engineering of an end-to-end system in which reachability, trust, operational authority, integrity, observability, and consequence remain deliberately bounded.
Secure connectivity: OT DMZs, jump hosts, ZTNA, and unidirectional flows
Industrial connectivity should begin from a simple architectural distinction: a requirement to exchange information or perform a remote function does not imply a requirement for general end-to-end network routability.
An industrial system may legitimately require remote maintenance, telemetry export, engineering access, enterprise reporting, cloud analytics, inter-site coordination, or communication with external service providers. None of those requirements implies that the participating networks should become mutually routable, that enterprise applications should acquire direct paths toward authoritative OT sources, or that successful authentication should automatically create broad network-level reachability inside the industrial environment.
This follows from both analytical models developed earlier. In the capability-inference model, connectivity supplies prerequisites for reachability-related derivations; in the IEC 62443 zone-and-conduit model, connectivity should realize the mission-required flow between security domains without silently creating additional relationships. The architectural objective is therefore to provide the specific communication required by the mission while minimizing the additional reachability, trust, and downstream authority created by providing it.
NCSC’s 2026 Secure connectivity principles for operational technology frames the problem similarly: connectivity should be justified by operational need, exposure should be limited, access paths should be centralized and standardized, boundaries should be hardened, compromise should be contained, connectivity should be monitored, and tested isolation capability should be maintained.48 The resulting design principle is:
Provide the required communication without converting that requirement into unnecessary network reachability or inherited trust.
The industrial DMZ should mediate trust, not merely route traffic
An industrial DMZ is not secure merely because it occupies a subnet between enterprise IT and OT. Its architectural purpose is to create a trust discontinuity: compromise or authority on one side should not automatically propagate through the DMZ into the other.
Where operationally feasible, cross-boundary communication should therefore terminate, be validated, and be re-originated by a service designed for the specific exchange rather than being forwarded transparently end to end. Typical DMZ functions include remote-access brokers, bastion services, controlled file-transfer staging, historian replication, protocol or application proxies, update staging, security-log aggregation, remote-session monitoring, and other services whose purpose is to exchange explicitly approved information or authority between trust domains.
The same architectural principle applies differently to data exchange and privileged administration. Figure 5 shows the two mediated patterns.
%%{init: {"theme": "neo", "look": "handDrawn", "layout": "elk"}}%%
flowchart TD
subgraph DATA["Mediated operational-data exchange"]
direction TB
A[Authoritative OT source]
B[/DMZ replica or<br/>application intermediary/]
C[Enterprise consumer]
A -->|approved data publication| B
B -->|controlled external consumption| C
end
subgraph ADMIN["Mediated privileged administration"]
direction TB
D[Trusted administrative endpoint<br/>or PAW]
E[/Access broker / bastion/]
F[[Mediated privileged session]]
G[Specific OT engineering resource]
D -->|authenticated administrative request| E
E ==>|policy grants bounded privilege| F
F -->|session restricted to approved target| G
end
Figure 5: Two mediated OT connectivity patterns. For data exchange, an OT source publishes through a DMZ replica rather than granting an enterprise consumer direct reachability to the authoritative OT source. For privileged administration, a trusted administrative endpoint reaches a specific OT resource through an access broker and mediated session. Rectangles denote ordinary endpoints or resources, slanted nodes denote mediation mechanisms, and double-bracket nodes denote privileged authority.
For enterprise consumption of OT data, the first pattern prevents the enterprise application from inheriting network reachability to the authoritative industrial source. A historian replica, message broker, application gateway, or other controlled intermediary can provide the required information while preserving the trust boundary. Direct enterprise queries into the OT environment may occasionally be justified, but they create a stronger dependency and therefore require stronger justification.
Privileged access follows the same architectural logic. A trusted endpoint should not ordinarily acquire general membership in the OT network simply because a user needs to administer one resource. Authentication, policy evaluation, session mediation, credential handling, target restriction, logging, and revocation should be concentrated at an explicit boundary wherever the process and technology allow it.
NCSC’s secure-connectivity guidance likewise recommends centralizing third-party remote access through hardened solutions associated with the DMZ rather than accumulating independent vendor VPN endpoints throughout the OT environment.49
This does not mean that every industrial flow must be proxied or terminated at the application layer. Some communications must remain direct because of protocol behavior, availability requirements, latency, deterministic operation, or vendor constraints. The narrower principle is that trust on one side of the DMZ should not propagate automatically to the other side.
A DMZ that merely forwards traffic after authentication while preserving broad downstream network reachability has changed the topology without necessarily changing the capability-inference structure.
Jump hosts matter only when they create a trust discontinuity
A jump host or bastion can centralize high-risk administrative access and enforce controls that cannot be applied directly to legacy industrial systems. Properly designed, it can separate external identities from downstream credentials, restrict available administration tools and destinations, mediate remote-desktop or terminal protocols, constrain file transfer, prohibit ordinary productivity workloads, record privileged sessions, and provide a controlled point for monitoring, activation, and revocation.
The same concentration that makes a bastion useful also makes it dangerous. A system through which many privileged relationships pass can become a high-consequence authority concentrator, particularly when it stores reusable downstream credentials or remains dependent on an identity environment that an attacker can compromise independently.
The Polish large-CHP case illustrates the broader risk of intermediary systems whose trusted credentials and administration functions remain reachable from an already compromised environment.50 A bastion that is broadly domain-joined, supports ordinary email or browsing, accepts connections from low-trust networks, and retains persistent credentials for many controllers may reduce neither attacker reachability nor derived authority; it can instead provide the adversary with a prepared concentration of the exact capabilities required for lateral administration.
A meaningful bastion must therefore create a genuine discontinuity in at least some of the prerequisites required for privileged access. It should constrain identity, endpoint trust, destinations, credentials, available tools, session duration, file movement, and permitted administrative operations according to the actual maintenance mission.
The relevant property is not the presence of a product labelled jump host. It is whether compromise on the upstream side remains insufficient to derive downstream administrative authority without crossing additional independently enforced conditions.
PAWs make the trustworthiness of the administrative endpoint explicit
A Privileged Access Workstation addresses a related but distinct problem. The bastion constrains the intermediate administrative path; the PAW constrains the trustworthiness of the endpoint from which privileged activity originates.
NCSC describes PAWs as trusted physical user devices intended to protect high-risk administrative access and discusses their application in OT environments. Its secure-connectivity guidance recommends restricting administration to PAWs where practicable, particularly for critical security controls and obsolete systems.51
For high-consequence OT administration, the originating endpoint should therefore be treated as part of the privileged-control architecture rather than as an ordinary enterprise workstation. It should expose substantially less attack surface and provide substantially fewer opportunities for unrelated compromise. Depending on the environment, this can justify restrictions on web browsing, email, arbitrary software installation, unmanaged removable media, uncontrolled file transfer, communication with unrelated enterprise systems, and use for non-administrative workloads.
The architectural principle is sometimes described as browse down: the device used to administer a system should be at least as trustworthy as the system or security domain being administered. The phrase is useful because it highlights the asymmetry of privileged control: strong downstream segmentation provides limited protection if the endpoint legitimately permitted to cross that segmentation has already been compromised.
This is particularly important in OT because an engineering endpoint can combine vendor software, controller projects, certificates, privileged credentials, trusted routes, and protocol capabilities that would otherwise have to be acquired separately. Compromise of such a workstation can therefore satisfy several prerequisites in a high-consequence derivation at once.
A PAW directly addresses the engineering-workstation failure identified in the earlier taxonomy, but it does not replace network segmentation, access mediation, identity controls, downstream authorization, or controller-native protections. It strengthens one part of the trust chain by reducing the probability that privileged activity begins from an already compromised endpoint.
ZTNA constrains admission, not industrial semantics
Zero Trust Network Access addresses another layer of the connectivity problem. NCSC’s 2026 ZTNA guidance treats access as an explicit policy decision based on identity and contextual signals rather than as an implicit consequence of network location. Its model combines policy decision and enforcement functions with signals describing users, endpoints, and access context, so that admission can be granted to a specific resource rather than to an entire routed network.52
That distinction can be expressed using a simple resource-admission model. Let \mathcal{R}_{\mathrm{routed}} denote the resources that are technically reachable from the network location presented to a remote user, let \mathcal{R}_{\mathrm{ZTNA}}(p,z) denote the resources admitted by policy for principal p under access context z, and let \mathcal{R}_{\mathrm{mission}}(p,z) denote the resources that principal actually requires for the approved operational task.
The first condition captures the difference from broad network extension: policy exposes only a subset of what the underlying routed environment could technically reach. The second is the least-functionality objective: the admitted resource set should coincide with the resources required by the approved task rather than with every resource available behind the access boundary.
NIST SP 1800-45 provides practical remote-access reference architectures for the water and wastewater sector using commercially available mechanisms for mediated access, authentication, authorization, and monitoring.53 Although sector-specific, the architectural principle is general: remote access does not require unrestricted network extension.
ZTNA nevertheless solves only part of the OT authorization problem. Suppose an engineer is legitimately admitted to an S7 controller. The ZTNA policy may establish that the identity, endpoint, time, source context, and target resource satisfy the admission policy, yet the industrial session established after admission may still permit diagnostic reads, project upload, project download, controller-mode changes, firmware replacement, or modification of protection-relevant parameters. Those operations are not equivalent simply because they traverse the same authenticated connection.
Within the capability-inference model, ZTNA primarily constrains the prerequisites associated with reachability and admission. It does not by itself determine the industrial semantics of every operation that becomes available after the session has been established. That requires downstream use control, engineering-system security, protocol-aware enforcement, controller-native authorization, maintenance-state restrictions, procedural controls, or some combination of them.
ZTNA therefore improves resource admission control, but it does not automatically provide industrial semantic authorization. A resource can be correctly admitted while still exposing a set of operations broader than the approved mission requires. This distinction is critical in OT. A security architecture is incomplete if it asks only whether a principal may connect to a controller; it must also determine which process-relevant capabilities become available after that connection succeeds.
Private transport is not a trust boundary
The Polish private-APN incident makes an architectural distinction explicit: private transport and endpoint trust are different security properties. A private APN, MPLS service, site-to-site VPN, carrier SD-WAN, or comparable service can prevent traffic from traversing the public internet and thereby reduce one important form of exposure, but that property does not establish the identity or integrity of connected endpoints, guarantee peer isolation, enforce least-privilege routing, provide administrative independence, or authorize the industrial operations that become possible once communication has been established.
Carrier infrastructure should therefore be treated as a conduit between security domains, not as evidence that every connected site belongs to one trusted network. The relevant architectural question is not whether two sites share private transport, but whether the industrial mission requires them to communicate and, if so, which specific flows are justified.
Using the mission-flow model introduced earlier, let \mathcal{F}_{ij}^{\mathrm{mission}} denote the communication flows required between remote sites S_i and S_j, and let \mathcal{F}_{ij}^{\mathrm{allowed}} denote the flows permitted by the deployed carrier and site architecture. When the industrial mission requires no communication between the two sites, the design requirement is
Membership in the same carrier service should therefore not create lateral reachability between sites that have no operational reason to communicate. Depending on the architecture, this may require site-specific routing, explicit peer isolation, ACLs or firewall policy, separation of carrier and plant management traffic, endpoint or gateway authentication, independent administration of the plant-side boundary, and monitoring of inter-site flows.
This directly addresses the shared-infrastructure failure identified earlier. A common communications service can provide valuable operational efficiency without becoming a common propagation mechanism, but only if the architecture deliberately prevents compromise of one participant from inheriting unnecessary reachability to the others.
Microsegmentation should follow consequence and trust
Microsegmentation can reduce lateral propagation inside an OT environment by applying policy at a finer granularity than conventional perimeter segmentation. NCSC’s secure-connectivity guidance explicitly identifies microsegmentation as one mechanism for limiting lateral movement and containing compromise.54 Granularity itself, however, is not a security objective: an architecture containing thousands of device-to-device rules can be highly restrictive on paper while becoming difficult to understand, test, maintain, and audit, and excessive rule complexity can itself create configuration drift and operational fragility.
Microsegmentation should therefore refine the zone-and-conduit architecture rather than replace architectural reasoning with individual firewall rules. Useful segmentation boundaries normally correspond to process functions, consequence classes, security zones, engineering relationships, trust boundaries, or classes of operational authority. Separating two process cells is meaningful when compromise of one must not acquire reachability or authority over the other; separating them merely because their devices occupy different address ranges provides no equivalent architectural justification.
The appropriate granularity should therefore be derived from the capability-inference model. The design should identify which capabilities must remain underivable after compromise of a process cell, engineering environment, remote site, or administrative domain, and then place segmentation where it removes the reachability prerequisites required by those derivations. This also means that different flows involving the same devices can require different treatment: read-only monitoring, operator control, engineering access, firmware management, and safety-related communication do not represent equivalent authority merely because they share endpoints. The relevant question is consequently:
Which capability derivations must remain impossible after compromise of this process cell or administrative domain, and which segmentation boundaries make their reachability prerequisites unavailable?
Connectivity must be designed to be withdrawn safely
Secure connectivity also requires a credible means of reducing or removing communication during an incident. NCSC’s secure-connectivity principles therefore include establishing and testing an isolation plan.55 In OT, however, an isolation plan cannot be limited to knowledge of which firewall rule, router interface, VPN, radio link, or physical cable must be disabled, because withdrawing connectivity can simultaneously remove legitimate sensing, supervision, engineering, synchronization, or control capability.
The requirement connects directly to the mission-viability model developed earlier. If an isolation action changes the cyber-operational state from z to z', the surviving control actions and realizable trusted control trajectories can change with it. Where isolation is intended to preserve the industrial mission, the resulting physical state must remain inside the mission-viability set \mathcal{K}_{\mathcal{M}}(z'); where continued operation is impossible, the isolation procedure must instead be coordinated with a validated transition to an approved degraded or shutdown state.
For every important external or inter-zone connection, the architecture should therefore establish how the connection can be isolated, who has authority to perform the action, how rapidly isolation can occur, which monitoring and control capabilities disappear afterward, which local or autonomous functions remain available, whether independent protection remains effective, and how the connection can later be restored without reintroducing attacker persistence or untrusted state. These properties should be exercised before an incident rather than discovered while containment is already under way.
Isolation is consequently another instance of the principle established earlier:
A cybersecurity action is acceptable in OT only when the resulting cyber-physical state and its effect on the physical mission are understood.
Data diodes eliminate the reverse path through a specific conduit
A hardware-enforced unidirectional gateway provides an unusually strong connectivity property because directionality is enforced physically rather than solely through configurable policy.
For a diode-controlled conduit d that exports information from OT toward an external domain, define the permitted directional-flow property as
The second condition is the defining security property: the conduit provides no reverse communication flow from the external domain toward OT. Unlike an ordinary firewall policy, that property cannot be reversed merely by changing a rule or compromising the management interface controlling the filter.
Within the capability-inference model, a diode can therefore remove an entire class of reachability prerequisites associated with that specific reverse conduit. NCSC’s secure-connectivity guidance likewise describes data diodes as mechanisms for physically enforced directionality while emphasizing that unidirectionality alone is not equivalent to a complete cross-domain security solution.56
The scope of the formal property is important. It applies to conduit d, not automatically to the complete system. A diode does not prove that exported information is trustworthy, that applications processing the information on the receiving side are secure, that content flowing in an architecturally permitted direction is semantically safe, or that another reverse path does not exist elsewhere through remote maintenance, management infrastructure, another network, or removable media.
Data diodes are consequently well suited to genuinely asymmetric functions such as historian replication, telemetry export, security-log export, selected monitoring feeds, and transfer of information toward isolated analysis environments. They are not appropriate where the industrial mission requires bidirectional interaction through the same relationship, nor should a bidirectional requirement be treated as solved merely by installing two opposite-direction diodes without analysing the application semantics reconstructed above them. Physical directionality constrains transport; it does not automatically make a higher-layer transaction safe.
The patterns solve different parts of the same problem
OT connectivity mechanisms should not be treated as interchangeable secure-access products because they constrain different parts of the trust and authority structure. A PAW strengthens the originating endpoint, a ZTNA or access broker constrains admission, an industrial DMZ mediates trust between domains, a bastion constrains privileged sessions, microsegmentation reduces lateral reachability, carrier peer isolation prevents unwanted inter-site propagation, and a data diode physically removes one direction of communication through a specific conduit.
Figure 6 shows how these mechanisms can be composed without implying that the topology is mandatory.
%%{init: {"theme": "neo", "look": "handDrawn", "layout": "elk"}}%%
flowchart TD
subgraph ADMIN["Privileged administration"]
direction TB
A[Named engineer<br/>trusted PAW]
B[/ZTNA / access broker/]
C[/Industrial DMZ<br/>session mediation/]
D[[Bounded privileged<br/>engineering session]]
E[Specific OT<br/>engineering resource]
A -->|authenticated administrative request| B
B -->|resource-specific admission| C
C ==>|mediated privilege granted| D
D -->|session restricted to<br/>approved target| E
end
subgraph DATA["Operational-data export"]
direction TB
F[OT historian]
G[/Constrained or unidirectional<br/>export mechanism/]
H[DMZ data replica]
I[Enterprise analytics]
F -->|approved operational data| G
G -->|controlled export| H
H -->|enterprise consumption| I
end
Figure 6: Reference OT connectivity pattern showing separate trust functions for privileged administration and operational-data export. Rectangles represent ordinary endpoints, resources, or reachable domains; slanted nodes represent mediation mechanisms; double-bracket nodes represent privileged authority. Thick arrows identify transitions in which bounded privileged authority is granted, while ordinary arrows represent permitted communication or progression.
The figure separates several security functions that are often collapsed into a single VPN or firewall rule: endpoint trust, identity, resource admission, trust mediation, privileged-session control, downstream target restriction, and data export. Collapsing them can be operationally convenient, but it also allows one compromised mechanism to satisfy several prerequisites in a high-consequence derivation simultaneously.
Table 7 summarizes the principal distinction among the patterns.
Pattern
Primary security property
Principal limitation
Flat network-extension VPN
Encrypts and authenticates a remote network connection
Often grants routability substantially broader than the resource or task actually required
Brokered remote access through an industrial DMZ
Centralizes admission, policy enforcement, session mediation, monitoring, activation, and revocation
The broker and its administration become high-value dependencies that require independent protection and resilience
Jump host / bastion
Constrains destinations, tools, credentials, file movement, and privileged sessions
Can become a credential and authority bridge if its own trust relationships are weak
PAW
Provides a high-trust originating endpoint for privileged administration
Does not by itself constrain downstream reachability or industrial operations
ZTNA
Makes resource admission identity- and context-dependent rather than network-location-dependent
Does not normally authorize the industrial semantics of the admitted protocol session
Microsegmentation
Restricts lateral reachability at finer granularity within or between zones
Excessive policy complexity can create governance, maintainability, and availability problems
Private APN / MPLS / carrier VPN
Removes or reduces dependence on public transport infrastructure
Transport privacy does not establish endpoint trust, peer isolation, least privilege, or operational authorization
Data diode / unidirectional gateway
Physically eliminates the reverse communication path through a defined conduit
Supports only genuinely asymmetric flows and does not constitute a complete cross-domain security solution
Table 7: OT connectivity mechanisms constrain different parts of the reachability, trust, and authority structure. They should be selected according to the communication and operational authority required by the mission rather than treated as interchangeable secure-access technologies.
The central design principle is therefore broader than segmentation. It is controlled composition of trust and authority. A remote engineer may need one engineering application without requiring membership in the plant network; an enterprise analytics platform may require process data without requiring a query path into OT; two remote sites may use the same carrier without requiring mutual reachability; an engineering workstation may need programming authority over one PLC without requiring persistent authority over every controller in the zone.
The secure architecture should preserve these distinctions so that fulfilment of one legitimate communication requirement does not silently satisfy the prerequisites for unrelated high-consequence capabilities. In terms of the high-consequence separation model introduced earlier, DMZ mediation, PAWs, bastions, ZTNA, carrier peer isolation, microsegmentation, and unidirectional gateways are valuable only insofar as the resulting architecture prevents the modeled initial footholds from deriving states in H.
None of these mechanisms is sufficient in isolation because each constrains a different part of the capability-inference system. The relevant connectivity question is therefore not:
How can this external system connect to OT?
It is:
What exact communication does the industrial mission require, what additional reachability or authority would that communication otherwise make derivable, and which independent controls ensure that the resulting relationship cannot be composed into a higher-consequence capability?
Hardening PLCs, engineering workstations, and industrial protocols
The preceding connectivity analysis primarily addressed whether a principal or system can establish a path toward an OT resource. Hardening addresses the next problem: what operational authority becomes technically available after that path already exists.
This distinction is fundamental because network reachability and controller authority are not equivalent. A historian can legitimately reach a controller while requiring only read access; an HMI may require a defined set of process commands without requiring firmware or project modification; an engineer may require programming authority over one controller during an approved intervention without requiring permanent authority over every controller in the process zone. A secure architecture therefore does not need to make every controller unreachable; it needs to prevent legitimate reachability from implying unnecessary operational authority.
For principal p, controller operating condition \mu, and relevant operational context z, let \Gamma_{\mathrm{allowed}}(p,\mu,z) denote the set of controller operations technically available to that principal, while \Gamma_{\mathrm{mission}}(p,\mu,z) denotes the subset of operations required to perform the approved industrial role.
The role must first be operationally feasible, so every mission-required operation must be available:
Define the excess operational-authority set inline as \Delta\Gamma(p,\mu,z)=\Gamma_{\mathrm{allowed}}(p,\mu,z)\setminus\Gamma_{\mathrm{mission}}(p,\mu,z). Under the mission-feasibility condition in Equation 30, least authority is therefore equivalent to
\Delta\Gamma(p,\mu,z)
=
\varnothing.
\tag{32}
Together, Equation 31 and Equation 32 give controller hardening a precise objective: the architecture must preserve every operation required by the approved industrial role while eliminating every additional operation that becomes technically available merely because the principal can reach and authenticate to the controller. The problem is therefore not to enable as many security features as the controller supports, but to eliminate the difference between what a reachable principal can technically do and what that principal must be able to do for the approved mission.
The distinction also prevents network security from being mistaken for operational authorization. A conduit may legitimately permit communication to a controller and authentication may correctly establish the identity of the engineer, yet the resulting session can still expose project download, firmware replacement, CPU mode change, credential administration, or unrestricted process-write functions that the current task does not require. Hardening must therefore continue beyond reachability and authentication into privilege, service exposure, controller configuration, engineering-system trust, and protocol semantics.
Disable capability that the mission does not require
The August Siemens advisory recommends reducing unnecessary controller exposure by disabling unneeded services and management functions, restricting communications, configuring available protection mechanisms, and limiting engineering access.57 These measures implement the same least-functionality principle at the controller and service layer.
For controller c, let \Sigma_{\mathrm{enabled}}(c) denote the services and management capabilities enabled in the deployed configuration, while \Sigma_{\mathrm{required}}(c) denotes those required by the industrial function and its approved maintenance model.
Unused FTP, Telnet, SSH, HTTP, SNMP, Modbus services, vendor diagnostics, programming interfaces, web-management functions, or firmware-management mechanisms should therefore not remain enabled merely because the product supports them. The exact list is device- and process-specific: some functions that appear unnecessary during normal operation may still be required for validated maintenance, recovery, commissioning, or emergency procedures and should be controlled accordingly rather than removed without analysis.
Service minimization is more than conventional attack-surface reduction because an exposed industrial management service can provide a prerequisite for substantially stronger capabilities. A reachable web interface may expose credential administration; an SSH service may permit tunnelling or configuration access; a programming interface may expose controller project operations; a diagnostic protocol may provide enough state information to enable later manipulation. The security significance of an enabled service therefore depends on which additional capabilities become derivable when that service is reachable and successfully used.
The Polish incidents demonstrate the importance of this distinction. Ordinary embedded management functions combined with weak or default credentials were sufficient to provide destructive administrative authority without requiring sophisticated exploitation of the underlying industrial device.58 Controller hardening should therefore remove or constrain unnecessary management capability before it can become part of a high-consequence derivation.
Network allowlists are useful, but network origin is not identity
Where supported, controller-side IP or MAC restrictions can reduce the set of systems permitted to communicate with an industrial device and can therefore provide useful compensating protection, particularly where legacy equipment cannot support stronger authentication or fine-grained authorization mechanisms.59 Their function must nevertheless be described correctly: network-origin restriction constrains reachability; it does not establish cryptographic identity or prove which human, process, or application is exercising the resulting authority.
An IP address can identify an expected network location while saying little about the actor currently controlling the host at that address, and a MAC address can restrict communication within a local network without providing a strong identity for the user or software issuing an industrial command. The distinction becomes especially important where several engineers share the same engineering workstation, applications execute under common service accounts, address translation obscures the originating endpoint, or an otherwise trusted host has itself been compromised.
Source allowlisting should therefore be treated as one layer of the reachability architecture rather than as a substitute for Identification and Authentication Control or Use Control. It can make communication from unexpected locations infeasible, thereby eliminating some prerequisites in the capability-inference model, but once a permitted endpoint has been compromised the allowlist may continue to admit exactly the traffic that the attacker requires. The engineering question is consequently not merely:
Is this source address permitted to reach the controller?
It is:
Which endpoint, principal, and application are exercising authority through the permitted network relationship, and which additional controls prevent compromise of that permitted source from becoming unrestricted controller authority?
Programming authority should be exceptional
Programming a controller is qualitatively different from reading process data or issuing an ordinary operator command because it can modify the logic that determines future process behavior rather than merely request an action within the logic already deployed. Firmware replacement, CPU operating-mode changes, protection-setting changes, security-configuration changes, and other functions capable of redefining controller behavior deserve similar treatment even when the vendor interface categorizes them as ordinary device administration.
Programming authority should therefore normally be contextual and exceptional rather than permanently available. The authorization decision should incorporate the identity of the engineer, the trustworthiness of the engineering endpoint, the specific target controller, the approved change activity, the controller’s operating or protection state, and the time interval during which the intervention has been authorized.
Let W_{\mathrm{prog}}(c) denote the approved programming windows for controller c. Where the technology and operating model permit time-bounded programming, a necessary condition for a successful programming operation is
The condition is deliberately necessary rather than sufficient. Being inside an approved maintenance window should not by itself grant programming authority; the action should additionally be attributable to a named engineer, originate from a trusted engineering endpoint, target an explicitly authorized controller, correspond to an approved change record, occur under an appropriate controller protection state, and produce auditable engineering evidence.
The stronger policy is therefore contextual:
The fact that an engineer is capable of programming a controller does not imply that the engineer should be able to program that controller at every time, from every endpoint, against every target, or outside an approved change activity.
The Polish second-CHP incident illustrates why operating-mode functions belong in the same high-consequence category. Placing PLC CPUs into STOP can disrupt the physical mission without modifying the control project at all; operating-mode authority must therefore be governed as an industrial control capability, not dismissed as routine device administration.60
Controller projects are production executable state
An industrial controller project is not merely source code stored for engineering convenience. Depending on the platform, it can contain ladder logic, structured text, function blocks, initialized data, hardware configuration, communication relationships, alarm definitions, process limits, controller-protection settings, and safety-relevant logic or configuration. Once deployed, these artifacts determine part of the executable state through which the cyber system controls the physical process.
Controller projects should therefore be treated as production executable state and governed with the same seriousness as other production configuration whose unauthorized modification can alter process behavior. The relevant integrity objective is not simply that a project file exists, but that the organization can establish through an authoritative, product-appropriate verification mechanism that the state actually running on the controller corresponds to an independently approved engineering baseline.
That correspondence need not mean byte-for-byte equality. Vendor toolchains may compile, normalize, regenerate metadata, or represent controller state differently from the engineering source, so the verification mechanism must be defined according to the platform. Depending on the controller family, evidence can include vendor project-comparison functions, controller upload and comparison, project or block checksums, signed engineering artifacts, version-controlled repositories, firmware and hardware configuration comparison, and verification of controller protection state.
The important distinction is:
A backup exists is weaker than the running controller can be shown to correspond to an independently trusted and approved state.
The first establishes that some recoverable artifact is available; the second establishes a basis for reasoning about controller integrity. Failure to make that distinction is the controller-integrity failure identified in the earlier taxonomy.
A gold copy requires provenance
A file becomes a trustworthy recovery baseline because its provenance, approval, validation, and relationship to the commissioned physical system are known; it does not become trustworthy because someone has labelled it a gold copy. An authoritative baseline should therefore identify, as applicable, the controller model and hardware revision, approved firmware version, engineering-project version, relevant integrity identifiers, approval history, safety-related configuration, communications configuration, required certificates or credential material, validation date and result, and the relationship between the digital project and the physical configuration that was commissioned.
The most recent project discovered on an engineering workstation is consequently not necessarily a valid recovery baseline, and the highest version number is not necessarily the approved production version. An engineer may have experimental, partially commissioned, obsolete, or locally modified projects that are technically loadable but operationally incorrect.
Recovery integrity therefore requires the organization to know which controller state is authoritative, which evidence establishes that authority, and why the baseline remained trustworthy while the production environment was potentially compromised. This also means that the recovery copy and the evidence supporting it should not depend exclusively on the same engineering environment whose compromise created the need for recovery.
Modern S7 security illustrates native capability versus architectural compensation
Siemens advisory SSA-568427 provides a useful example of the transition from legacy trust assumptions toward stronger native security mechanisms. The advisory explains that legacy protection across affected SIMATIC S7-1200 and S7-1500 families relied on a built-in global private key whose protection was no longer considered sufficient, while TIA Portal V17 and corresponding firmware generations introduced device-specific protection of confidential configuration data and TLS-protected PG/PC and HMI communication. For supported configurations, Siemens recommends migrating projects and enabling secure PG/PC and HMI communication.61
The broader architectural lesson is not specific to Siemens. Where the controller, engineering environment, HMI, and other required peers all support a stronger secure communication mode, compatibility with obsolete insecure modes should not silently preserve the weaker trust model unless a documented operational constraint requires it. Native security capability should be enabled when it can materially remove authentication, integrity, or confidentiality weaknesses without violating process requirements.
Native protocol security and system architecture must nevertheless remain distinct. A modern controller may provide stronger cryptographic identity and communications protection, while a legacy controller may not; the absence of a native capability does not remove the underlying security requirement but shifts more of the burden toward compensating architecture. Stronger zone boundaries, dedicated engineering workstations, narrow conduits, source allowlisting, controlled maintenance windows, mediated access, passive monitoring, and physical restrictions can reduce the authority that remains derivable even when the endpoint itself cannot implement the preferred control.
This is the same distinction established earlier between component capability and achieved system security. Native protection can make a stronger architecture possible, but its mere presence does not prove that the complete system has been configured so that unacceptable authority remains unavailable.
The engineering workstation belongs to the controller’s trusted computing base
Engineering workstations contain the legitimate capabilities required to redefine industrial behavior. They can combine vendor programming suites, authenticated controller relationships, certificates and credentials, production projects, firmware tools, device-discovery functions, privileged routes, removable-media access, and the ability to perform operations that ordinary operator stations cannot execute. A compromised engineering workstation can therefore satisfy several prerequisites for high-consequence controller authority simultaneously, providing an attacker with capabilities that would otherwise have to be acquired through separate technical steps.
For high-consequence environments, engineering endpoints should consequently be treated as privileged infrastructure rather than as ordinary enterprise workstations. Appropriate restrictions can include eliminating routine email and unrestricted web browsing, tightly constraining internet connectivity, using dedicated privileged identities, enforcing application allowlisting, limiting host-network policy to required destinations, controlling removable media and file transfer, deploying endpoint protection that has been validated for the engineering toolchain, protecting project and credential storage, and recording privileged engineering actions.
The architectural objective is not merely to make the workstation difficult to compromise. It is to reduce the number of unrelated compromise mechanisms that can provide access to the same endpoint from which legitimate controller authority originates. An engineering workstation that combines ordinary productivity use, internet browsing, enterprise authentication, vendor tooling, persistent controller credentials, and privileged OT routes collapses several trust domains into one endpoint and thereby concentrates consequence.
AA26-097A demonstrates another important aspect of this problem: legitimate engineering applications can themselves be used to perform malicious industrial actions.62 The August Siemens advisory separately warns about snap7-based scripts represented as legitimate monitoring utilities.63 These cases show why executable identity and engineering-action legitimacy are different properties.
Application allowlisting can prevent execution of software that is not approved, but it cannot establish that an approved programming suite is being used for an approved engineering purpose. A legitimate executable can perform an illegitimate project download, mode change, firmware operation, or parameter modification if the surrounding authorization model permits it.
The stronger monitoring and authorization object is therefore the engineering action in context, including the actor, originating endpoint, target controller, operation, change record, time, and resulting controller state, rather than merely the hash or publisher of the executable that generated the traffic.
Industrial protocol security has three distinct layers
Industrial protocol security is often discussed as though encryption or TLS were sufficient to make an industrial protocol secure. That collapses three different questions: whether the peer is authentic, whether the message is protected against relevant forms of tampering or disclosure, and whether the industrial operation carried by that valid message is actually authorized.
For industrial message m, define the acceptance predicate
Here, A_{\mathrm{peer}}(m) represents the required authentication and trust of the communicating peer; I_{\mathrm{message}}(m) represents the integrity, freshness, and, where the application requires it, confidentiality properties of the message; and Z_{\mathrm{operation}}(m) represents authorization of the requested industrial operation for the current identity, role, target, controller state, and operational context.
The conjunction is analytically useful because the three predicates are not interchangeable. Peer authentication can establish who or what originated a request, message integrity can establish that the protected request was not modified in transit, and encryption can conceal its contents from unauthorized observers, yet none of those properties determines whether the requested operation is appropriate.
A cryptographically authentic, integrity-protected S7 STOP operation can still be an unauthorized S7 STOP operation. Likewise, an authenticated Modbus client can still issue a write outside its operational role, and a correctly authenticated engineering application can still download an unapproved project.
This residual distinction between secure communication and legitimate industrial action is the protocol-semantic failure identified in the taxonomy.
Modbus Security
Traditional Modbus/TCP does not itself provide modern cryptographic peer authentication or message-integrity protection suitable for operation across an untrusted network. Modbus Security combines the Modbus application protocol with TLS, uses X.509v3 certificates for mutual client/server authentication, provides message-integrity protection, and uses TCP port 802. It can also convey role-based authorization information through certificate extensions, although the policy by which those roles are interpreted and enforced remains dependent on the implementation.64
Modbus Security therefore materially strengthens peer authentication and protected message transport, but those improvements do not imply that every authenticated client should be authorized to perform every Modbus operation. The residual architectural question remains semantic: which authenticated principal may read which data objects, which principals may write, which write operations are legitimate in the current process context, and how is that authorization enforced by the endpoint or surrounding architecture?
A deployment that adds TLS while preserving unrestricted write authority for every authenticated engineering client has solved an important communications-security problem without solving the complete industrial authorization problem.
CIP Security
ODVA’s CIP Security provides security mechanisms for CIP and EtherNet/IP using established cryptographic protocols. For EtherNet/IP, TLS protects applicable TCP-based communication while DTLS protects applicable UDP-based communication; CIP Security supports device authentication using X.509 certificates or pre-shared keys, message integrity, optional confidentiality, and security profiles that have expanded to include user-level authentication and authorization capabilities.65
This is materially richer than treating CIP Security as transport encryption alone because the available mechanisms can address peer identity, protected communication, and, where the appropriate profiles are supported and configured, user-level authorization. The achieved security property nevertheless depends on which profiles the devices actually implement, which profiles have been enabled, how certificates or pre-shared keys are governed, how users and roles are managed, which permissions are assigned to authenticated identities, and whether the authenticated endpoints themselves remain trustworthy.
The last point is important. Cryptographic device identity reduces impersonation and message tampering; it does not prove that the authenticated device has not been compromised. A legitimately authenticated engineering station or controller can remain a source of malicious industrial traffic if an attacker has acquired control of that endpoint and can exercise its existing authority.
CIP Security should therefore be evaluated as part of the complete authority model rather than as an endpoint property considered in isolation.
DNP3 security
DNP3 security mechanisms similarly address more than simple encryption. The DNP Users Group’s Cybersecurity Task Force develops and maintains DNP3 Secure Authentication and the associated Authorization Management Protocol, with current work addressing authentication, integrity, authorization, key management, and related security controls.66
These mechanisms strengthen the system’s ability to establish whether sensitive DNP3 operations originate from authorized entities and whether protected messages retain their required integrity, but they do not eliminate the consequences of compromise of an already trusted master station, operator identity, engineering environment, or authorization infrastructure. Once a principal has legitimately acquired an authorization context, the architecture must still constrain the industrial actions available through it according to operational role and process consequence. The residual question is therefore:
Is this authenticated and integrity-protected DNP3 operation appropriate for this principal, this target, and the current process context?
Secure Authentication is necessary where the threat model requires it, but it cannot substitute for governance of the authority that successfully authenticated principals are allowed to exercise.
OPC UA
OPC UA provides a particularly rich native security model among the protocol families considered here. Its architecture distinguishes application authentication, user authentication, SecureChannel protection, message signing, encryption, certificate trust, user authorization, role-based permissions, access restrictions, and security auditing.67 These mechanisms allow security policy to be expressed at a finer level than in protocols whose security capabilities concentrate primarily on transport protection or device authentication.
That richness should not be confused with secure deployment. The specification provides mechanisms that system designers and product implementations can configure according to the installation, but it does not prove that a particular deployment has disabled insecure modes, established sound certificate lifecycle management, configured appropriate TrustLists, assigned roles correctly, constrained permissions to mission requirements, or enabled auditing capable of reconstructing security-relevant activity. The distinction again mirrors IEC 62443:
Security capability is not the same as achieved security.
OPC UA can provide strong primitives for authenticated, integrity-protected, confidential, and role-constrained communication, yet those primitives become an effective security property only when certificate governance, application trust, user identity, role assignment, address-space permissions, auditing, endpoint hardening, and network architecture compose correctly.
Table 8 compares the protocol families according to the security properties they can provide and the residual authority that must still be governed.
Protocol family
Security mechanism
Principal security properties
Residual architectural question
Siemens S7 engineering
Secure PG/PC and HMI communication on supported modern generations
TLS-protected communication and stronger protection of confidential configuration data
Which authenticated engineering operations are actually permitted, and how are legacy communication modes constrained?
Modbus
Modbus Security
TLS, X.509 mutual authentication, message integrity, and support for role information
Which authenticated clients may perform which reads and writes, against which targets and under which operational conditions?
EtherNet/IP / CIP
CIP Security
TLS/DTLS, device authentication, integrity, optional confidentiality, and available user-authentication and authorization profiles
Which profiles are implemented and enabled, how are identities governed, and are endpoint, user, and role privileges appropriately bounded?
DNP3
Secure Authentication and authorization-management mechanisms
Authentication, integrity, authorization, key management, and related protection for sensitive operations
What authority remains available if an already authorized master, operator identity, or authorization path is compromised?
OPC UA
SecureChannels, certificates, users, roles, permissions, access restrictions, and auditing
Application and user authentication, integrity, confidentiality, authorization, and auditability
Are secure modes, TrustLists, certificates, roles, permissions, endpoint trust, and auditing correctly governed in the deployed system?
Table 8: Industrial protocol security mechanisms constrain different parts of the authority structure. Cryptographic protection reduces impersonation and tampering, but protected transport does not by itself establish that an authenticated industrial operation is operationally legitimate.
Industrial authority is the composition of several security gates
No single hardening mechanism answers the complete industrial authorization problem because different controls constrain different prerequisites for high-consequence authority. Endpoint trust addresses whether the workstation from which privileged activity originates can itself be trusted; the conduit determines whether that endpoint can communicate with the target at all; protocol security establishes peer and message properties; controller use control determines which functions an admitted identity may exercise; change governance determines whether a high-consequence operation is permitted in the current context; project-integrity verification determines whether resulting controller state corresponds to an approved baseline.
These controls should therefore be understood as independent security conditions surrounding a capability progression, rather than as interchangeable layers through which an attacker or administrator merely passes.
Figure 7 represents that relationship using the same capability, condition, authority, and consequence semantics as the earlier diagrams.
%%{init: {"theme": "neo", "look": "handDrawn", "layout": "elk"}}%%
flowchart TD
A[Named engineer]
B[Trusted engineering<br/>workstation]
C[Specific controller<br/>reachability]
D[[Programming / write<br/>authority]]
E{{Controller logic,<br/>configuration or operating state}}
F(((Physical process<br/>effect)))
G{Authorized conduit<br/>and target restriction}
H{Authenticated peer and<br/>protected message}
I{Controller identity and<br/>use-control policy}
J{Approved change context<br/>and programming window}
K{Verified project provenance<br/>and approved baseline}
A -->|uses privileged endpoint| B
B -->|mission-required communication| C
G -.->|constrains permitted reachability| C
H -.->|constrains accepted communication| D
C ==>|reachable resource can expose<br/>privileged operations| D
I -.->|bounds permitted functions| D
J -.->|bounds high-consequence change| D
D ==>|authorized operation modifies<br/>controller state| E
K -.->|constrains acceptable deployed state| E
E ==>|control state changes<br/>physical behavior| F
Figure 7: Composition of security conditions governing high-consequence industrial authority. Rectangles represent ordinary actors, endpoints, reachability, or deployed resources; braces represent architectural or authorization conditions; double-bracket nodes represent privileged operational authority; double-curly nodes represent high-consequence controller or process-control state; the terminal rounded node represents physical consequence. Dotted arrows identify contributory security conditions, ordinary arrows represent use of an existing capability, and thick arrows represent escalation of authority or consequence.
The Figure 7 deliberately separates capability progression from the conditions that constrain it. The engineer and workstation do not acquire programming authority merely because a communication path exists; the thick transition from controller reachability to programming or write authority represents an escalation whose feasibility depends on identity, use control, protected communication, target restriction, and approved change context. Likewise, possession of programming authority does not establish that the resulting deployed state is legitimate: project provenance and baseline verification constrain whether a controller modification should be accepted as trustworthy.
Each condition therefore answers a distinct engineering question. Endpoint trust asks whether the system exercising authority is itself sufficiently trustworthy; conduit policy asks whether that endpoint should reach this target; protocol protection asks whether the peer and message can be authenticated and protected; controller use control asks which functions the admitted identity may exercise; change governance asks whether a high-consequence operation is permitted now and for this intervention; project-integrity verification asks whether the resulting controller state corresponds to an independently approved baseline; process analysis asks what physical consequence that controller state can produce.
The capability-inference model developed earlier gives this composition its formal interpretation without requiring additional symbolic arrow expressions. A well-designed architecture makes high-consequence authority dependent on several independently enforced prerequisites, so compromise of one layer does not automatically satisfy the conditions required by the next. Reachability alone should therefore remain insufficient for programming authority, authentication alone should remain insufficient for arbitrary command authority, and possession of an approved engineering tool should remain insufficient for unrestricted process modification.
This is why controller hardening, engineering-workstation security, secure industrial protocols, contextual change authorization, and project-integrity verification are complementary rather than interchangeable. Segmentation determines who can arrive at the controller; hardening and authorization determine what remains possible once they arrive; integrity verification determines whether the resulting controller state can still be trusted. Their composition determines whether an otherwise ordinary digital compromise can accumulate into high-consequence industrial authority and ultimately alter the physical process.
Detection in deterministic industrial networks
Prevention is necessarily incomplete. Credentials will sometimes be stolen, trusted endpoints will sometimes be compromised, approved engineering software can be abused, and a controller may accept a syntactically correct, authenticated, and even formally authorized command that is malicious in the operational context in which it is issued. Detection therefore addresses a different problem from prevention: it must determine whether the observed exercise of industrial authority remains consistent with the architecture, operating context, controller state, and physical behavior that should exist.
OT provides defenders with an important advantage because many industrial communication relationships and process behaviors are comparatively stable, repetitive, and engineered in advance. NIST notes that OT traffic is generally more deterministic—more repeatable and predictable, than conventional enterprise traffic, which can support anomaly and error detection.68Deterministic should not be interpreted as perfectly periodic or immutable: legitimate behavior changes with production state, startup, shutdown, maintenance, failover, commissioning, and process conditions. The useful property is narrower and stronger for defensive purposes: the set of legitimate peers, protocols, operations, roles, and process behaviors is often substantially more constrained and explainable than in a general-purpose enterprise network.
Detection can exploit that structure by treating the intended industrial architecture not only as a preventive design but also as a measurable baseline against which the deployed system can be continuously compared.
Architecture can become the detection baseline
The zone-and-conduit architecture developed earlier specifies which communication relationships are intended to exist, while operational context determines which subset of those relationships should be active at a particular time. The same architectural representation can therefore support detection if expected communication is made explicit rather than left implicit in network diagrams and firewall configurations.
Let z denote the current operational context, such as normal production, startup, shutdown, maintenance, commissioning, or another condition that legitimately changes network behavior. Represent the expected communication structure under that context as G_0(z)=\left(N_0(z),E_0(z)\right), and the communication structure observed during monitoring interval t as G_t=\left(N_t,E_t\right).
Here, A_E^{+}(t;z) contains communication relationships that are present but not expected, A_E^{-}(t;z) contains expected relationships that have disappeared, and A_N^{+}(t;z) contains newly observed nodes that do not belong to the expected context-specific architecture.
All three classes can be security-relevant. A new enterprise-to-PLC connection, an HMI initiating engineering traffic, or one remote station communicating laterally with another can indicate that an unauthorized path has appeared; disappearance of a persistent controller-to-controller exchange, historian feed, protection-system relationship, or supervisory connection can indicate equipment failure, deliberate disruption, segmentation failure, or attacker activity. An unexpected node can similarly represent temporary maintenance equipment, configuration drift, an unmanaged engineering station, or an unauthorized device.
Detection therefore becomes a form of continuous architectural validation. The operational question is no longer merely whether individual packets appear malicious, but whether the network currently behaves like the zone-and-conduit architecture that the organization claims to have engineered.
The August Siemens-focused advisory recommends hunting for indicators including unexpected S7 communication from systems that should not perform engineering functions, scanning of TCP/102, anomalous data access, suspicious writes, snap7-related activity, and engineering operations inconsistent with expected behavior.69 The specific indicators vary by technology, but the architectural principle is general: unexpected relationships have unusual evidentiary value in OT because legitimate relationships are comparatively constrained.
Passive observation should be preferred where active probing creates process risk
Asset discovery and network monitoring must themselves respect the cyber-physical constraints established earlier. Active scanning, malformed probes, aggressive service discovery, or unexpected industrial-protocol sequences can destabilize some legacy or fragile equipment, so a detection architecture should not create process risk merely to measure cybersecurity state. NIST therefore recommends understanding the operational impact of active techniques and identifies passive monitoring of normal network traffic as an important means of learning OT communications and distinguishing known from unknown behavior.70
SPAN ports, physical network taps, switch telemetry, controller logs, identity records, engineering-host events, configuration baselines, and process telemetry can provide complementary evidence without requiring the monitoring platform to initiate control traffic. Figure 8 represents this design principle.
%%{init: {"theme": "neo", "look": "handDrawn", "layout": "elk"}}%%
flowchart TD
A[Remote / enterprise systems]
B[Industrial DMZ]
C[Engineering environment]
D[Process cells]
E[/Passive protocol-aware<br/>network observation/]
F[Identity and access<br/>evidence]
G[Engineering and controller<br/>evidence]
H[Approved project and<br/>configuration baselines]
I[Process telemetry and<br/>independent measurements]
J[OT security analytics]
K[SOC / incident analysis]
A --> B
B --> C
C --> D
B -.->|TAP / SPAN evidence| E
C -.->|TAP / SPAN evidence| E
D -.->|TAP / SPAN evidence| E
E -.->|network evidence| J
F -.->|identity context| J
G -.->|engineering and state evidence| J
H -.->|trusted comparison baseline| J
I -.->|physical evidence| J
J -->|correlated findings| K
Figure 8: Layered OT detection architecture in which passive network observation and independent telemetry provide evidence without becoming prerequisites for real-time process control. Rectangles represent ordinary systems or evidence sources and slanted nodes represent observational mechanisms. Ordinary arrows represent the underlying permitted OT connectivity, while dotted arrows denote observational evidence paths rather than process-control dependencies.
Where feasible, the monitoring architecture should remain observational rather than operationally prerequisite. Failure of a passive sensor should reduce visibility without preventing a PLC from continuing its control function, because unnecessary coupling between detection and control would introduce the monitoring system into the same failure model that cybersecurity is supposed to improve.
This does not make monitoring availability unimportant. Loss of visibility is itself a security degradation because it reduces the organization’s ability to detect misuse of authority, validate system state, and reconstruct events. The narrower architectural principle is that loss of a detection component should not automatically become loss of the physical mission unless that dependency is explicitly required and engineered as such.
Baseline relationships, not just traffic volume
A useful OT baseline describes the semantics of communication relationships, not merely average bandwidth, packet rate, or flow duration. An attacker can preserve aggregate traffic volume while radically changing the industrial meaning of the traffic, so a controller that continues to exchange approximately the same number of packets may nevertheless have moved from read-only supervision to project transfer or destructive mode control.
For device v under operational context z, let B(v,z) denote its context-sensitive communication baseline. The baseline should describe at least the expected peer set, communication directionality, permitted protocols and industrial operations, expected timing or activity patterns, and the architectural role of the device.
For a PLC, that baseline might state that a historian may read process values, an HMI may issue a defined set of operational commands and reads, a peer controller may participate in a fixed cyclic exchange, an engineering workstation may perform infrequent engineering operations only during approved activities, and an enterprise workstation has no direct relationship at all.
This is substantially stronger than saying that the PLC normally transfers a certain number of megabits per second. The useful monitoring object is the relationship, role, and operation, because those properties determine whether observed communication is consistent with the authority model established by the architecture.
Protocol-aware detection should decode industrial operations
The hardening section distinguished peer authentication, message protection, and authorization of the industrial operation. Detection adds a separate question: even if the communication is accepted by the receiving system, is the observed operation expected in the current operational context?
Represent a semantic industrial event as e=(s,d,p,o,a,t), where s is the source, d the destination, p the industrial protocol, o the decoded industrial operation, a the relevant operation arguments, and t the event time.
For operational context z, define the expectation predicate
This model separates four distinct conditions: whether the communicating peers should have a relationship, whether the protocol is permitted on that relationship, whether the specific industrial operation and its arguments are appropriate, and whether the event is consistent with the current operating context.
It supports detection rules such as a historian issuing a write, an HMI attempting a project transfer, an unexpected workstation initiating PLC programming, a vendor session changing CPU mode, an engineering operation occurring outside an approved maintenance activity, a normally read-only relationship suddenly performing writes, or one remote industrial site initiating traffic toward another where no mission requirement exists.
The distinction from protocol acceptance is essential. A controller can accept a command because its peer is authenticated, the message is intact, and the access-control policy permits the operation, while the same command remains anomalous because it is inconsistent with the current mission, maintenance state, engineering authorization, or process condition. Compromise of a legitimate engineer account is precisely the kind of scenario in which protocol validity and operational expectation diverge.
Maintenance context belongs inside the detection model
OT environments legitimately behave differently during commissioning, maintenance, testing, startup, shutdown, and emergency intervention. A PLC project download that is highly anomalous during normal production may be expected during a planned outage, while a new engineering relationship that would normally trigger immediate investigation may be explicitly authorized for a short maintenance activity. Detection therefore requires operational context as part of the model rather than as an explanation appended after an alert has fired.
Useful contextual evidence can include the work order or change record, remote-access approval, named engineer, PAW identity, broker session, target controller, industrial operation, approved project version, maintenance window, and plant operating mode. These attributes should contribute directly to z, because the meaning of an observed event depends on whether the surrounding engineering and process state makes that event legitimate.
NCSC notes that monitoring rules may sometimes be adjusted or temporarily suppressed during maintenance to reduce false positives, while emphasizing that the SOC should remain informed, normal monitoring should be restored outside planned windows, and the claimed maintenance activity must itself be verified.71 The correct conclusion is therefore not:
Disable OT detection during maintenance.
It is:
Make maintenance a first-class detection context while preserving evidence about what occurred during the maintenance activity.
An adversary who has compromised a maintenance account should not become invisible merely by operating during a declared maintenance window. Maintenance changes what is expected; it should not eliminate accountability.
Monitor controller state as well as network traffic
Network monitoring can establish that a command crossed a conduit, but it does not necessarily establish what state the controller subsequently entered. Controller integrity therefore requires device-state evidence in addition to network evidence, particularly for changes such as RUN/STOP transitions, operating-mode changes, firmware updates, project or block modifications, protection-level changes, identity or credential changes, communications-configuration changes, device resets, or firmware replacement.
Let S_{\mathrm{controller}}(t) denote the observed security-relevant controller state at time t, and let S_{\mathrm{approved}}(t;z) denote the state approved for the current operational context. Because controller state contains heterogeneous categorical and structured attributes, subtraction or ordinary numerical distance is not generally meaningful; the relevant question is whether the deployed state corresponds to the independently approved state under a product-appropriate verification relation.
Here, \equiv_{\mathrm{verified}} is the verification relation introduced earlier for project and controller integrity. A value M_S(t;z)=1 means that at least one security-relevant element of deployed controller state cannot be shown to correspond to the approved state for the current context.
That condition is not automatically evidence of attack. It may result from an approved change that has not yet been reconciled with the baseline, maintenance activity, equipment replacement, configuration drift, operator error, or malicious manipulation. It is nevertheless a condition that requires explanation because the architecture can no longer demonstrate correspondence between expected and deployed controller state.
The monitoring objective is therefore not simply to detect packets or commands. It is to determine whether network activity, engineering actions, and resulting controller state remain mutually consistent with approved operational activity.
Network and controller evidence remain predominantly digital, which means that a sufficiently privileged attacker may be able to manipulate several of those evidence sources simultaneously. The physical process can provide another evidentiary layer when independent or partially independent measurements allow the organization to compare reported digital state with physical behavior.
Consider the simplified discrete-time balance for a tank:
where V_k is tank volume at sample k, q_{\mathrm{in},k} is measured inflow, and q_{\mathrm{out},k} is measured outflow. If the reported level remains constant while sufficiently independent measurements indicate sustained net inflow, the digital representation and physical conservation model are inconsistent.
More generally, let \hat{y}_k be the output predicted by an appropriate process model and y_k the observed output. With residual r_k=y_k-\hat{y}_k, a context-sensitive anomaly detector can evaluate
\left\|
r_k
\right\|
>
\theta(z).
\tag{41}
where \theta(z) is a threshold appropriate to the operating context and model uncertainty.
A large residual indicates that observed process behavior is inconsistent with the model to the degree specified by the detector; it does not identify the cause. Equipment failure, sensor degradation, calibration error, changed feedstock, unmodeled disturbances, process transitions, and modeling error can produce the same effect as malicious manipulation.
Process anomaly is therefore evidence of inconsistency, not attribution of cyberattack. Its value lies in providing an evidence channel that can be at least partially independent of the digital control path the attacker may be manipulating, which directly addresses the process-observability failure identified in the earlier taxonomy.
Cross-layer correlation is stronger than any single signal
The strongest OT detection rarely comes from one event considered in isolation. Identity authentication, remote-session establishment, execution of an engineering tool, an industrial programming connection, a controller-state transition, and a change in physical output can all occur legitimately; their security significance emerges from their composition, ordering, common target, and operational context.
Figure 9 illustrates such a sequence using the same capability, authority, and consequence semantics as the earlier diagrams.
%%{init: {"theme": "neo", "look": "handDrawn", "layout": "elk"}}%%
flowchart TD
A[Vendor identity<br/>authentication]
B[Brokered remote<br/>session]
C[Engineering tool<br/>to controller reachability]
D[[Programming or operating-mode<br/>authority]]
E{{PLC CPU<br/>STOP state}}
F(((Turbine power<br/>reduction)))
A -->|authenticated access| B
B -->|mediated engineering path| C
C ==>|industrial privilege exercised| D
D ==>|mode-changing operation| E
E ==>|controller state affects process| F
Figure 9: Cross-layer evidence sequence in which legitimate-looking identity and access events progress into privileged industrial authority, a high-consequence controller state, and physical consequence. Rectangles denote ordinary events or reachable technical states, double brackets denote privileged authority, double-curly nodes denote high-consequence controller state, and the terminal rounded node denotes physical consequence. Thick arrows identify escalation of authority or consequence.
Outside an approved maintenance or incident-response context, this composition is substantially stronger evidence than any one event considered independently. Identity records establish who or what authenticated; broker logs establish through which privileged path access occurred; host evidence establishes which trusted software executed; network and protocol evidence establish which industrial operations crossed the conduit; controller evidence establishes which state actually changed; and process evidence establishes what happened physically.
Correlation should nevertheless distinguish multiple observations from independent evidence. Two monitoring tools are not independent if both derive controller state from the same compromised PLC, and several dashboards do not provide independent physical confirmation if they all consume the same manipulated process tag. Evidence becomes particularly valuable when compromise of one trust path cannot silently redefine every observation layer.
Controller-state evidence, packet captures, identity records, engineering logs, independent sensors, and physical measurements should therefore be correlated while preserving knowledge of their provenance and failure domains.
Break-glass access should be intentionally noisy
Emergency access creates exceptional authority precisely because ordinary security constraints are being bypassed or relaxed. It should therefore be straightforward to invoke when genuinely necessary but difficult to use silently, and its activation should produce immediate, deterministic security evidence rather than depend on statistical anomaly scoring.
NCSC states that any attempt to use a break-glass account should trigger the highest-criticality alarm within the SOC.72 The same principle applies to other exceptional operations such as disabling protection, changing safety-related parameters, replacing firmware, creating new privileged identities, or bypassing normal access mediation.
For these events, rarity and consequence are already known by design. Detection should therefore be explicit: every invocation should create an attributable, high-priority event and preserve sufficient evidence for immediate review. Behavioral analytics can provide additional context, but they should not determine whether an event that is intrinsically exceptional deserves attention.
Evidence should survive destruction of the device it describes
Detection is useful only if the resulting evidence survives long enough to support containment, reconstruction, and recovery. The Polish follow-up illustrates how destructive actions and emergency restoration can remove or overwrite local forensic evidence, making evidence location and independence part of the monitoring architecture rather than an afterthought.73
High-value evidence should therefore not exist only on the component whose compromise, reset, reimaging, or destruction it is intended to explain. The exact replication mechanism depends on the architecture, but useful patterns include remote syslog, OT-local collectors, security appliances, replicated event repositories, append-resistant or immutable storage, passive packet capture, controller-management platforms, and tightly constrained export toward a SOC or forensic repository.
The important property is failure-domain separation. Destruction of a controller should not automatically destroy the only evidence of its engineering changes; compromise of an engineering workstation should not automatically permit alteration of the only record of privileged activity; rebuilding an HMI should not erase the only timeline of the network relationships that preceded the incident.
This connects directly to the secure-connectivity architecture developed earlier because evidence export is often naturally asymmetric. Logs, events, packet metadata, or selected telemetry may need to move from OT toward a security repository without creating a corresponding administrative path from that repository back into control. Where operationally appropriate, tightly constrained or physically unidirectional export can preserve that distinction.
Table 9 summarizes the resulting evidence architecture.
Evidence layer
Core question
High-value examples
Identity and access
Who entered the privileged path, using which identity, endpoint, and authorization context?
MFA, vendor identity, unusual source, dormant account, break-glass use
Endpoint and engineering
Which trusted tool executed, and which engineering activity was initiated?
New peers, unexpected zone crossing, scanning, lateral remote-site communication
Industrial protocol
Which industrial operation was requested or executed?
Writes, project transfer, CPU mode changes, firmware operations
Controller integrity
Which security-relevant state did the controller enter?
RUN/STOP, project checksum, firmware, identities, protection level, IP or configuration changes
Physical process
Does observed physical behavior remain consistent with commands, models, and sufficiently independent measurements?
Flow, pressure, level, temperature, power, and command-response inconsistencies
Table 9: Six complementary evidence layers for OT detection. Their value comes from correlation across distinct trust domains and failure modes rather than from dependence on one telemetry source.
Detection can therefore be understood as the measurement system of the security architecture. Preventive design establishes which identities, communication relationships, industrial operations, controller states, and process consequences should remain possible under each operational context; detection measures whether the deployed system continues to satisfy those assumptions and provides evidence when observed behavior diverges from them.
Prevention and detection are complementary rather than alternative strategies. Prevention attempts to make dangerous capability derivations infeasible, while detection observes the relationships and states that remain possible, identifies deviations from the expected architecture, and correlates digital activity with controller and physical consequence. The strongest OT detection architecture therefore does not ask merely:
Does this packet resemble an attack?
It asks:
Is the observed identity, communication relationship, industrial operation, controller state, and physical response consistent with the architecture and operational context that should exist, and, if not, which capability or trust assumption appears to have failed?
Isolation, backup, recovery, and safe reconstitution
Incident response in OT cannot be reduced to quarantine the endpoint. Containment, recovery, and restart are themselves cyber-physical operations because disconnecting a system can remove attacker reachability while simultaneously removing legitimate control capability; restoring a controller can recover automation while destroying forensic evidence; reloading an independently approved project can restore digital integrity while leaving the controller’s internal representation inconsistent with the physical state that evolved during the outage.
Recovery must therefore solve several coupled problems. Active adversary authority has to be contained, the physical process must remain safe or reach an approved safe state, trustworthy digital state must be re-established, recovered automation must be reconciled with physical reality, mission-essential control capability must be restored, and external connectivity must be reintroduced without reconstructing the authority path that enabled the incident. These activities have dependencies, but they do not constitute a universal linear runbook: evidence capture can overlap containment, local fallback can operate while digital recovery proceeds, and the exact order depends on the process, incident scope, architecture, and independently available safety mechanisms.
The governing principle is that restoration of digital services is not equivalent to restoration of the industrial mission. NIST’s 2026 OT Backup Quick Start Guide emphasizes integrating OT backup management with change management, creating backups at appropriate intervals, testing them, and validating recovery through exercises.74 NCSC similarly calls for tested OT isolation plans and recommends that critical functions be designed, where possible, to operate independently of external dependencies.75
Isolation is a vector, not a switch
An OT incident rarely presents only two useful states, connected and disconnected. Containment can instead be applied independently at several layers: an identity can be revoked, a privileged session terminated, a remote-access service disabled, a conduit blocked, a security zone isolated, a remote site disconnected, or external connectivity removed more broadly. Treating all of these actions as one binary notion of isolation obscures both their different security effects and their different consequences for the physical mission.
Represent the containment posture by a vector \mathbf{I} whose components correspond to containment at the identity, session, service, conduit, zone, site, and external-connectivity layers. Applying a particular posture \mathbf{I} changes the cyber-operational state from z to z_{\mathbf{I}}, thereby changing both the inference rules available to an attacker and the trusted control capability available to legitimate operators.
For the incident foothold C_{\mathrm{inc}}, containment succeeds on the cyber side only if the resulting state removes derivability of the relevant high-consequence capabilities:
C_{z_{\mathbf{I}}}^{*}(C_{\mathrm{inc}})
\cap
H
=
\varnothing.
\tag{42}
That condition is necessary but not sufficient. The resulting cyber-operational state must also remain physically admissible: either the current physical state remains inside the mission-viability set associated with z_{\mathbf{I}}, or sufficient trustworthy control capability remains to move the process into an approved degraded or shutdown state without leaving the safe region.
Containment is therefore a constrained architectural intervention. The preferred response is not necessarily the largest isolation action that can be performed, but the narrowest independently trustworthy action that reliably removes the dangerous capability derivation while preserving the greatest amount of safe operational capability. Revoking one compromised vendor identity may be preferable to disconnecting an entire plant, disabling one remote-access conduit may be preferable to isolating every remote facility, yet narrow containment is inadequate when the integrity of the broader identity, management, or communications domain can no longer be established.
The containment mechanism must itself remain trustworthy
Containment has security value only if the mechanism enforcing it lies outside, or can be made independent of, the trust domain whose integrity is in doubt. If the firewall through which the adversary maintained access is also the sole mechanism expected to enforce emergency isolation, adding a new rule to that firewall may provide little assurance when the device itself, its administrative credentials, its configuration, or its management path may already be compromised.
The response architecture should therefore provide escalating containment options whose enforcement becomes progressively less dependent on the suspected failure domain. Depending on the system, responders may be able to revoke an identity, terminate a privileged session, disable a remote-access service, block a conduit, filter traffic at an independently administered upstream boundary, isolate a security zone, disconnect a site, or physically interrupt a communications path. The order is not mandatory and should not be interpreted as a universal escalation procedure; its purpose is to ensure that the organization has alternatives when the first containment mechanism cannot itself be trusted. The architectural requirement is:
The organization must retain at least one containment mechanism whose authority does not depend entirely on the component, identity plane, management plane, or trust domain suspected of compromise.
This is the incident-response analogue of the independence principle developed earlier for safety, monitoring, and recovery. A security boundary whose enforcement disappears when the protected domain is compromised is not an independent containment boundary.
External dependencies should be classified before the incident
Isolation cannot be designed safely if the organization does not know which external dependencies the physical mission requires. Enterprise identity services, cloud authentication or licensing, vendor-support platforms, carrier networks, remote engineering services, centralized historians, enterprise name resolution, centralized time services, inter-site control, cloud optimization, and external dispatch systems can all become operational dependencies even when they appear to reside outside the nominal control system.
For external dependency d, let z^{-d} denote the cyber-operational state in which that dependency is unavailable. The useful question is not merely whether an application continues to run under z^{-d}, but whether the physical mission remains viable and, if it does not, whether the process retains sufficient trustworthy authority to reach and remain in an approved degraded state.
Each material dependency should therefore be classified before an incident into at least three categories:
Mission-preserving: loss of the dependency still leaves the current physical state inside the mission-viability set, so the required mission can continue using surviving trusted capabilities.
Safe-degraded: the complete mission cannot continue, but the process can remain inside the safe region or transition into an approved degraded or shutdown state.
Mission-critical dependency: loss of the dependency removes control, observation, protection, or coordination capability required either to maintain the mission or to execute the required safe transition.
For the strongest case, mission preservation means that the current process state remains viable under the dependency-loss state:
This classification converts isolation planning from an infrastructure exercise into a cyber-physical resilience analysis. Discovering during an active compromise that Internet isolation disables the identity service required by local operators, or that loss of a carrier path also removes the only trusted time source required by protection functions, means that the dependency analysis was performed too late.
Manual and local fallback constrain consequence
Manual operation, local control, autonomous fallback, and islanded operation can materially reduce the consequence of a digital compromise because they preserve legitimate control capability after centralized or remote authority has been removed. The July 2026 FBI and EPA warning concerning attacks against U.S. water-sector PLCs explicitly recommends retaining the ability to operate OT manually and routinely testing business-continuity, fail-safe, islanding, backup, and standby mechanisms.76 The Polish CHP incident similarly demonstrates how timely operator intervention can constrain physical consequence after digital control has been disrupted.77
Manual fallback should not be interpreted as evidence that preventive cybersecurity succeeded. It is a resilience and consequence-containment mechanism that becomes valuable precisely when part of the digital security architecture has failed.
Let z_{\mathrm{degraded}} represent a degraded state in which centralized digital capability has been lost, and let z_{\mathrm{fallback}} represent the corresponding state after validated manual, local, or autonomous fallback has been made available. A successful fallback mechanism enlarges or preserves the set of trustworthy actions available to the plant:
The additional trusted authority can preserve mission viability or make safe shutdown possible even when remote engineering, supervisory control, or centralized optimization is unavailable. That benefit exists only when the fallback is credible: personnel must be trained, procedures must be accessible and current, local controls and instrumentation must function, process limits must be known, staffing and response times must be realistic, and the operating mode must be exercised periodically under representative degraded conditions.
A manual procedure that has not been executed for many years is therefore not equivalent to tested fallback capability, and manual operation is not inherently safe for every process. The architecture must determine which functions can genuinely be transferred to human or local control, what information those operators will still possess, how long the degraded mode can be sustained, and which safety protections remain independent of the failed digital path.
Safety constrains forensic preservation
OT incident response creates an unavoidable tension between forensic preservation and physical stabilization. Factory-resetting a controller can remove attacker persistence and permit rapid recovery while also destroying volatile logs or other evidence; preserving the device untouched may retain valuable forensic material while leaving the process exposed to continuing malicious authority or preventing restoration of an essential control function.
The Polish CHP recovery illustrates this tension because restoration activity was necessary to recover industrial operation while resets and reconstitution also reduced the forensic evidence remaining on affected devices.78
There is therefore no universal ordering in which forensics must always precede recovery or recovery must always precede evidence preservation. Process safety is the hard constraint. Incident-response actions must first belong to the set of actions that preserve the safe region or establish an approved safe state; only within that admissible response space should responders optimize among competing goals such as containment strength, service continuity, forensic preservation, recovery speed, and operational disruption. The practical principle is:
Preserve evidence wherever doing so does not conflict with process safety, necessary containment, or physical stabilization.
The architecture should compensate for this unavoidable trade-off by moving high-value evidence outside the components most likely to be reset, reimaged, replaced, or disconnected during an incident. Remote logs, passive network captures, independent controller-state records, engineering repositories, identity events, and external evidence stores are therefore not merely monitoring conveniences; they preserve investigative capability when the safest operational action requires destruction or replacement of the original source.
OT backup scope is broader than server backup
OT backup cannot be reduced to copying application files because recovery may require reconstruction of a distributed cyber-physical configuration spanning firmware, controller logic, device parameters, network state, identities, certificates, supervisory applications, field-device settings, and the documentation necessary to reconstruct their dependencies.
For asset a, represent the recoverable state bundle as \mathcal{B}(a)=(F_a,P_a,C_a,N_a,I_a,K_a,D_a), where F_a covers required firmware or software state, P_a the executable project or control application, C_a the device configuration, N_a the relevant network state, I_a identity and access-control configuration, K_a cryptographic material or the information required to restore or reissue it securely, and D_a the documentation, dependency information, and recovery metadata required to use the other elements correctly.
The concrete bundle varies by asset class. For a PLC it can include the approved project, hardware configuration, firmware reference, communications configuration, protection settings, user configuration, and safety-relevant engineering state. For a firewall or industrial security gateway it can include security policy, routing, NAT where applicable, VPN definitions, certificates, identity mappings, logging configuration, and management-plane configuration. For HMI or SCADA systems it can include the operating system or runtime requirements, application version, tag database, drivers, displays, scripts, certificates, user and role configuration, and service dependencies.
The word backup can therefore be misleading when interpreted only as stored files. The real requirement is recoverable authoritative state: the organization must possess enough trustworthy information, software, configuration, credentials, tooling, dependencies, and evidence to reconstruct the intended operational function rather than merely restore a collection of bytes.
Backup should follow engineering change
A recovery baseline becomes stale when authoritative production state changes, so backup management must be coupled directly to engineering change management. An approved material change to controller logic, firmware, communications configuration, identity policy, safety parameters, supervisory applications, or another recovery-relevant state should result in corresponding validation and update of the authoritative recovery baseline.
This relationship is more meaningful than imposing one arbitrary backup frequency on every OT asset. A controller project that changes twice per year may not require nightly export of identical project data, but it does require a verified recovery baseline after each approved material change; a rapidly changing supervisory database or recipe system may require substantially more frequent protection because its acceptable recovery point and rate of legitimate state change are different.
Backup cadence should therefore reflect the rate of legitimate change, the consequence of losing state, the acceptable recovery point, the feasibility of reconstructing the state from other authoritative sources, and the recovery objectives of the physical mission. NIST SP 1339’s emphasis on integrating backup management with change management is especially important in OT because the newest backup is useful only if it represents an approved and recoverable production state.79
Backup quality is multidimensional
The mere existence of a backup artifact proves little about recoverability. For recovery artifact b, distinguish five properties: validity V(b), meaning that the artifact is structurally usable and not corrupt; completeness C(b), meaning that it contains the state required for the intended recovery; freshness F(b), meaning that it satisfies the required recovery point; trust T(b), meaning that provenance and integrity are sufficiently established; and restorability R(b), meaning that the organization has demonstrated that the artifact can actually restore the intended function.
A high-quality recovery artifact must satisfy all five properties, which can be summarized as Q(b)=V(b)\land C(b)\land F(b)\land T(b)\land R(b). The conjunction matters because the properties fail independently: a backup can be valid but incomplete, complete but obsolete, current but already contaminated by attacker modification, trustworthy but unusable because the required engineering tool or license no longer exists, or perfectly restorable in a laboratory while failing the recovery-time constraints of the actual industrial mission.
Testing is therefore not an administrative formality but evidence for restorability. Recovery exercises also expose hidden dependencies that static inventories rarely capture, including obsolete engineering software, missing licenses, undocumented certificate chains, unavailable hardware revisions, inaccessible vendor tools, non-recoverable credentials, or process assumptions known only to individual engineers.
Backup quality should consequently be measured by demonstrated recoverability of the intended industrial function, not by backup-job success counters alone.
Recovery copies should not share production authority
A recovery repository does not provide resilience if the same compromised authority can alter production state and destroy or corrupt every trusted recovery copy. The problem is architectural rather than merely a question of storage configuration: backup systems administered through the same privileged identities, management workstations, directory infrastructure, or unrestricted administrative paths as production can become part of the same compromise domain.
Recovery architecture should therefore reduce high-authority dependencies shared with production. Appropriate mechanisms can include offline copies, immutable or write-once storage, independently administered repositories, separate backup identities, separation of backup credentials from ordinary production administration, approval requirements for destructive backup operations, and geographically or logically independent copies.
The objective is to prevent one credential compromise, privileged-workstation compromise, ransomware event, or management-plane compromise from becoming a common-mode production-and-recovery failure. This is the recovery analogue of the safety-independence principle developed earlier: redundancy has limited value when all copies remain subject to the same destructive authority.
Independence should also be evaluated against recovery operations themselves. A backup repository that is well isolated during normal production but can only be accessed through a compromised engineering workstation at the moment of restoration may still fail to provide a trustworthy reconstitution path.
A backup must not silently preserve compromise
Independence protects recovery material from destruction; it does not prove that the contents of the material are trustworthy. A perfectly immutable copy of malicious controller logic remains malicious controller logic, and a complete backup containing attacker-created identities, modified firewall rules, compromised certificates, or unauthorized firmware can reintroduce persistence during recovery.
Recovery must therefore distinguish an available backup from a trusted recovery baseline. Before restoration, responders may need to establish when compromise began, whether the candidate baseline predates that boundary, whether controller logic, firmware, boot state, privileged identities, certificates, network policy, or engineering tools were affected, and whether the environment used to validate the backup is itself trustworthy.
Where the compromise boundary cannot be established with sufficient confidence, recovery may require reconstruction from an earlier independently validated engineering baseline rather than restoration of the most recent copy. This is why backup management and controller-integrity management are the same architectural problem viewed at different points in time: one establishes the approved state before compromise, while the other uses that evidence to reconstruct a trustworthy state afterward.
Restoration should follow trust dependencies
Recovery order cannot be universal because industrial systems have different dependency structures, but one principle is general: a recovered component should not be reintroduced into a dependency that remains untrusted when doing so would immediately invalidate the trust established by recovery.
A clean engineering workstation restored into a compromised privileged-identity environment can immediately become exposed again; a verified controller project loaded through an untrusted engineering workstation does not establish controller integrity; a clean firewall configuration administered with compromised credentials does not reconstitute a trustworthy boundary; and a correctly restored PLC started against an unknown physical process state can be digitally correct while remaining physically unsafe.
Safe reconstitution should therefore follow trust and process dependencies, rather than restoring systems only according to application criticality. Figure 10 represents a conceptual dependency structure rather than a universal runbook.
%%{init: {"theme": "neo", "look": "handDrawn", "layout": "elk"}}%%
flowchart TD
A[Stabilize physical process<br/>and contain active authority]
B[Preserve evidence<br/>where operationally safe]
C[Establish trusted recovery<br/>environment and baselines]
D[Re-establish identity and<br/>administrative trust]
E[Re-establish trusted boundaries<br/>and required conduits]
F[Recover trusted<br/>engineering capability]
G[Verify controller firmware,<br/>configuration and protection state]
H[Restore independently approved<br/>controller projects]
I[Restore HMI / SCADA<br/>and supporting services]
J[Validate telemetry,<br/>alarms, monitoring and protection]
K[Reconcile digital and<br/>physical process state]
L[Controlled mission<br/>restoration]
M[Progressive external<br/>reconnection]
A -->|safe conditions for recovery| C
A -->|when operationally possible| B
B -->|preserved evidence supports validation| C
C -->|trusted recovery foundation| D
D -->|trusted administrative authority| E
E -->|trusted access paths| F
F -->|trusted engineering path| G
G -->|verified controller platform| H
H -->|approved controller state| I
I -->|supervisory capability restored| J
J -->|validated observation and protection| K
K -->|cyber-physical state aligned| L
L -->|mission proven before exposure| M
Figure 10: Conceptual OT safe-reconstitution dependency structure. Rectangles represent operational recovery states or activities and ordinary arrows represent prerequisite or progression relationships, not attacker-capability escalation. Physical stabilization and containment establish the conditions for trusted recovery; identity, boundaries, engineering systems, controllers, supervisory systems, and process state are then progressively validated before external connectivity is restored.
The diagram deliberately places progressive external reconnection after internal trust, process reconciliation, and mission restoration. A plant should not recreate the relationships and conditions that made the original high-consequence derivation feasible merely because individual components appear functional again.
Each external dependency should be reintroduced only after the architecture has established the intended identity, route, conduit, privileges, protocol behavior, monitoring, and operational need. Reconnection is therefore another security decision, not simply the final networking step of disaster recovery.
Physical and digital state must be reconciled
One of the most distinctive OT recovery problems occurs when digital systems have been restored but the physical plant has continued to evolve during the outage. A controller project can restore logic, configuration, internal variables, and setpoints, but it cannot automatically determine where every physical component is currently positioned or what operators changed while centralized automation was unavailable.
During degraded operation, personnel may have moved valves manually, opened or closed breakers, started or stopped pumps locally, changed bypasses, altered process inventory, isolated equipment, or changed controller modes. The state represented by the restored automation can therefore diverge from physical reality even when every restored digital artifact is independently trustworthy.
Let \hat{x}_r denote the recovered digital estimate of process state immediately before restart and let x_r denote the physical state established through sufficiently independent observation. Partition the safety-relevant state variables into continuous variables \mathcal{I}_c and discrete variables \mathcal{I}_d.
where \epsilon_i is the engineering tolerance defined for variable i. For every discrete state variable j\in\mathcal{I}_d, numerical subtraction is generally meaningless; reconciliation instead requires the recovered digital state to correspond to the independently verified physical state according to the state-specific equivalence criterion defined for that variable.
Reconciliation may therefore require independent confirmation of valve and breaker positions, reservoir or tank levels, pressures, temperatures, machine positions, rotating-equipment state, manual or automatic selector state, bypasses, interlocks, and process inventory. Which variables matter depends on the process, but the principle is general: a trustworthy digital controller state does not by itself establish that the physical initial conditions assumed by that controller are correct. A digitally clean controller can therefore produce an unsafe response if it resumes execution against the wrong physical state.
Restart is a controlled transition, not a binary event
After digital and physical state have been reconciled, restart should still not be treated as one command that returns the plant immediately to nominal operation. Where the process permits it, restoration should proceed through explicitly validated operating states such as safe shutdown, degraded controlled operation, partial restoration, mission-capable operation, and finally normal operation.
At each transition, operators should validate the properties that the next state assumes: process safety, controller state, alarms, protection functions, command-response behavior, communications, monitoring, identity and access state, and the health of newly restored dependencies. The exact sequence must be process-specific, because a chemical plant, substation, manufacturing line, water facility, and battery energy-storage system do not share one physically meaningful restart progression.
This staged approach is particularly important after destructive attacks because apparent digital recovery can conceal incorrect controller state, residual unauthorized configuration, missing telemetry, disabled protection, stale credentials, incomplete synchronization, or attacker persistence elsewhere in the architecture. Progressive restoration creates observation points at which these inconsistencies can be detected before the system reaches a state with greater physical consequence.
Mission restoration is complete only when the organization has sufficient evidence that the digital architecture and physical process are simultaneously operating within their approved states. Restart is therefore a controlled cyber-physical transition whose completion must be demonstrated, not inferred from server availability or successful controller boot.
Progressive reconnection should test whether the high-consequence derivation remains infeasible
External communications should be restored gradually rather than by recreating the pre-incident network architecture in one step. For each conduit being reintroduced, the organization should establish why the relationship is operationally required, which endpoints may use it, which flows are permitted, which identities and credentials remain trusted, which industrial operations can traverse the relationship, and which monitoring evidence will demonstrate that the conduit behaves as intended.
The recovery test is not merely that the associated service functions again. Reintroducing connectivity changes the capability-inference system, so the recovered architecture must continue to satisfy the same high-consequence separation requirement used during initial design:
C_{z_r}^{*}(C_0)
\cap
H
=
\varnothing
\qquad
\forall C_0\in\mathfrak{C}_0.
\tag{47}
If restoring a VPN, carrier route, vendor-access service, identity federation, cloud connection, or engineering conduit makes a previously blocked high-consequence derivation feasible again, then the system has restored digital availability while also restoring the vulnerability structure that enabled the incident.
Progressive reconnection should therefore be treated as controlled re-expansion of the trust and reachability architecture. Each restored relationship must earn its place through demonstrated operational need and verified containment properties rather than through a desire to return automatically to the pre-incident topology.
Recovery is restoration of trustworthy mission capability
Keep the process inside safe bounds or move it to an approved safe state
Cyber response itself produces or amplifies physical consequence
Cyber containment
Remove the adversary’s active high-consequence authority
The attacker remains able to manipulate systems while they are being recovered
Evidence preservation
Retain sufficient independent evidence where operationally safe
Root cause, incident scope, persistence, or trusted recovery point cannot be established
Trusted recovery environment
Establish independently trusted baselines, tooling, repositories, and recovery identities
Recovery artifacts or tools may already belong to the compromised trust domain
Identity and boundary recovery
Re-establish trusted administration, identities, zones, conduits, and enforcement mechanisms
Clean components are reintroduced into an untrusted control plane
Engineering and controller recovery
Restore verified firmware, configuration, projects, identities, and protection state
Controller state cannot be shown to correspond to approved engineering state
Supervisory recovery
Restore HMI, SCADA, communications, alarms, telemetry, and operator visibility
Operators lack sufficiently trustworthy observation or supervisory authority
Physical reconciliation
Align recovered digital assumptions with independently established physical state
Correct automation acts from incorrect physical initial conditions
Controlled restart
Validate process behavior through progressively more capable operating states
Hidden cyber or physical inconsistencies become immediate process consequence
Progressive reconnection
Restore only validated external dependencies and conduits
Recovery recreates the original reachability, trust, or authority path
Table 10: OT recovery is the reconstitution of trustworthy cyber-physical mission capability rather than restoration of computer uptime.
The objective is therefore stronger than restore the systems. Successful recovery requires active adversary authority to have been removed, critical digital components to have a defensible chain of trust, recovery artifacts and tooling to be independently trustworthy, physical process state to be known, controller assumptions to correspond to physical reality, mission-essential trusted control capability to be restored, and external connectivity to no longer recreate the capability derivations that enabled the incident.
Using the models developed earlier, safe reconstitution seeks a recovered cyber-operational state z_r and a reconciled physical state x_r such that
The first condition states that the physical mission is again viable using the trusted control capability available in the recovered architecture. The second states that restoring that capability has not simultaneously restored a derivation from the design-basis compromise scenarios to an unacceptable capability or cyber-physical state.
That conjunction is safe reconstitution. Its purpose is not merely to make destructive authority temporary; it is to make the industrial mission recoverable without making the original compromise reproducible.
Supply-chain security and secure-by-design industrial products
The preceding architecture assumes that the underlying products make its security properties technically realizable. An asset owner cannot create controller audit events when the controller emits none, cannot establish uniquely accountable identities when the product exposes only a shared password, cannot cryptographically verify firmware integrity when the device provides no trustworthy verification mechanism, cannot eliminate an insecure legacy protocol when that protocol remains mandatory for normal operation, and cannot continue patching indefinitely after the supplier has ended security support while the industrial asset remains operational.
Supply-chain security therefore concerns more than hardware provenance or the identity of the manufacturer. It concerns the lifecycle through which industrial capability, operational authority, trust, maintainability, and recoverability are created and sustained, including the software, firmware, engineering tools, cloud services, remote-support mechanisms, identity dependencies, update channels, third-party components, and supplier processes on which the deployed product continues to depend.
This connects directly to the IEC 62443 distinction developed earlier between component capability and achieved system security. For product p, let \mathcal{P}_{\mathrm{native}}(p) denote the security properties available natively from the product, let \mathcal{P}_{\mathrm{required}}(p) denote the properties required by the system architecture, and let \mathcal{P}_{\mathrm{comp}}(p;K) denote properties supplied through compensating architectural controls K.
The inclusion does not imply that every product deficiency can be compensated externally. Some security properties are inherently difficult, incomplete, or unsafe to recreate outside the device, including trustworthy boot, internal privilege separation, verification of running firmware, cryptographic validation of updates, protection of device secrets, and sufficiently detailed generation of controller-native audit evidence. Product selection therefore constrains the architecture that can subsequently be built: compensating controls can redistribute enforcement, but they cannot make arbitrary deficiencies disappear.
The 2024 international Principles of Operational Technology Cyber Security identifies supply-chain security as one of its foundational OT principles, while the 2025 international Secure by Demand guidance translates the same concern into acquisition criteria for owners and operators selecting digital products.8081
Secure by design moves responsibility upstream
Secure-by-design engineering changes where preventable security weakness is expected to be eliminated. When a weakness can be removed once during product design, requiring every asset owner, system integrator, and site operator to rediscover and compensate for the same weakness independently creates repeated implementation cost, inconsistent outcomes, and a fleet-wide opportunity for configuration error.
Universal default credentials are a clear example. If a product ships with the same reusable secret across every installation, security depends on every deployment correctly replacing that credential during commissioning and maintaining the replacement throughout the lifecycle. Some deployments will inevitably retain the insecure state, particularly where commissioning is rushed, documentation is incomplete, ownership changes, or devices remain untouched for years. Eliminating universal defaults at the product boundary removes that recurring failure mode before deployment rather than transferring it to every operator.
CISA and the FBI identify universal default passwords as a dangerous product-security practice and similarly warn against supplying products used in critical infrastructure with known exploitable vulnerabilities under the conditions described in their product-security guidance.82
The broader principle is therefore:
Do not repeatedly transfer to every operator a security problem that can be removed once at the product boundary.
This does not remove operator responsibility. Products still require risk-specific configuration, integration, maintenance, and monitoring, but secure-by-design practice reduces the number of predictable weaknesses that operators must first undo before meaningful system security can begin.
Secure by default is stronger than supports security
A product can support strong security mechanisms while remaining insecure in its ordinary deployed state. That distinction is especially important in OT because commissioning configurations can survive for years or decades, so functionality that exists but must be discovered, licensed, manually enabled, or carefully reconstructed after installation may never become part of the actual protection architecture. The relevant acquisition question is therefore not only:
Can this security feature be enabled?
It is:
What security state does the product create when a competent integrator follows its ordinary supported installation and commissioning path?
Secure by Demand recommends products that provide secure baseline configuration, elimination of default passwords, contemporary secure protocols, baseline logging, strong authentication, configuration management, secure communications, vulnerability-management capability, usable upgrade mechanisms, and meaningful operator ownership.83
For product p, the residual hardening gap can be represented by the set difference \Delta_{\mathrm{hardening}}(p)=\mathcal{P}_{\mathrm{required}}(p)\setminus\mathcal{P}_{\mathrm{factory}}(p), where \mathcal{P}_{\mathrm{factory}}(p) contains the security properties actually active in the supported baseline configuration. Secure-by-default design seeks to minimize that gap and, for properties that can reasonably be enabled universally, make it empty.
This does not mean that every deployment should share one universal configuration. Risk-specific segmentation, identities, privileges, trust anchors, logging destinations, maintenance policies, and protocol relationships remain necessary. The objective is narrower: a product should not require the operator to remove avoidable and well-understood insecurity merely to reach a defensible starting point.
Backward compatibility can preserve obsolete trust
Protocol migration illustrates another product-level difficulty. A controller may support a modern secure communication mode while continuing to accept a legacy mode for compatibility with older engineering tools, HMIs, gateways, or peer controllers. In that configuration, availability of the secure protocol does not eliminate the security properties of the legacy one; if the weaker relationship remains accepted, an attacker or compromised trusted endpoint may still be able to use it.
The problem is especially acute when negotiation can silently fall back to weaker operation or when secure mode is offered as an optional enhancement while the legacy interface remains permanently exposed. In such cases, the product has added cryptographic capability without necessarily changing the effective trust model of the deployed system. The procurement question should therefore not stop at:
Does this device support TLS, certificates, secure S7 communication, CIP Security, OPC UA security, or another secure protocol?
It should continue with:
Can obsolete compatibility modes be disabled once every mission-required peer supports the secure mode, and can secure-only operation be verified as part of the deployed configuration?
Backward compatibility is sometimes operationally necessary, particularly during staged brownfield migration, but it should be treated as an explicit architectural exception with a defined dependency and migration plan rather than as an invisible permanent feature.
Configuration state should be governable
The detection and recovery sections established that the operator must be able to determine what state is deployed, compare that state with an approved baseline, and reconstruct it after failure or compromise. Product capability must make those activities technically possible.
A governable industrial product should therefore provide, through mechanisms appropriate to its technology, the ability to export or otherwise retain authoritative configuration, identify the deployed version or state, verify relevant integrity properties, compare deployed state with an approved baseline, and restore the intended configuration. A PLC may expose project checksums or vendor comparison tools; a firewall may provide configuration export, revision history, and integrity mechanisms; an HMI may require application packages, runtime configuration, certificate stores, scripts, drivers, and supporting software state to be managed together.
The architectural requirement is not that every product expose an identical interface. It is that production state can be identified, independently retained, compared with approved state, and restored using supported mechanisms whose results can be evidenced.
Without that property, configuration drift becomes difficult to distinguish from compromise, incident investigation becomes dependent on assumptions about what the device should contain, and recovery becomes dependent on the availability of individual engineers or undocumented vendor procedures rather than on authoritative technical evidence.
Logging should be a baseline product capability
The detection architecture developed earlier depends on evidence, and some of that evidence can only be generated reliably by the product performing the action. A device capable of high-consequence administration but incapable of exposing meaningful security events creates an observability gap that passive network monitoring cannot always repair because network capture may show that a session occurred without proving which authenticated user acted, which internal privilege was exercised, or which persistent state changed.
Where technically appropriate, industrial products should therefore expose events such as successful and failed authentication, privileged-session establishment, project upload or download, logic modification, CPU or controller-mode changes, firmware updates, protection-setting changes, identity and role administration, communication-configuration changes, resets, security-policy modifications, and safety-relevant administrative operations. The exact event model is necessarily product-specific, but the principle is not:
Security-relevant actions should generate usable evidence in the baseline product rather than requiring an unrelated premium analytics function merely to establish that the action occurred.
This does not mean that sophisticated analytics must be included in every controller. The distinction is between generating primary security evidence and performing higher-level correlation: the former belongs close to the action being performed, while the latter can be supplied by the monitoring architecture developed earlier.
Ownership means operational autonomy over mission-critical functions
Industrial ownership is weaker than physical possession of hardware when mission-critical operation, maintenance, recovery, or administration remains dependent on supplier-controlled services that the asset owner cannot independently govern. Supplier cloud platforms, license-validation infrastructure, proprietary engineering repositories, remote-maintenance gateways, certificate services, update infrastructure, vendor identity systems, and diagnostic platforms can all become part of the effective operating boundary of an otherwise on-premises product.
Secure by Demand explicitly treats owner autonomy as an acquisition consideration and asks whether operators can maintain, modify, recover, and support products without unnecessary continuing dependence on the supplier.84
For product p, let \mathcal{D}_{\mathrm{mission}}(p) denote the external dependencies whose unavailability or compromise can materially affect a mission-relevant product function. These dependencies belong to the cyber-physical failure model even though they may be geographically remote or commercially outside the traditional equipment boundary.
The mission-viability framework developed earlier provides the appropriate test. For dependency d\in\mathcal{D}_{\mathrm{mission}}(p), the operator should determine whether loss of d preserves mission viability, forces a safe degraded state, or removes trusted capability required to keep the process safe. A cloud licensing service, remote identity provider, vendor engineering portal, or proprietary certificate infrastructure can therefore be operationally relevant even while the PLC and process network remain physically local.
The August 2026 U.S. bulk-power executive order illustrates the same widening of the supply-chain boundary at national-security scale by addressing associated software, firmware, digital services, maintenance services, remote-access capabilities, lifecycle mechanisms, and other dependencies rather than treating bulk-power equipment as isolated physical hardware.85
Secure product development is a lifecycle property
Product security cannot be inferred from the absence of a publicly known vulnerability at one point in time. A product with no currently disclosed CVEs can result from an immature development and disclosure process, while a supplier that regularly publishes and remediates vulnerabilities may be demonstrating that its vulnerability-handling lifecycle is functioning rather than proving that its products are uniquely insecure.
IEC 62443-4-1 addresses secure product development across security requirements, secure design, implementation, verification and validation, defect management, patch management, vulnerability handling, and end-of-life processes.86 NIST’s Secure Software Development Framework provides a complementary general model and vocabulary that acquirers can use when communicating secure-development expectations to suppliers.87
The procurement question is therefore not merely:
Does the current firmware have known CVEs?
It is:
Does the supplier operate a repeatable lifecycle capable of preventing, discovering, communicating, correcting, distributing, and eventually retiring security-relevant product defects?
For long-lived industrial equipment, this lifecycle property matters more than the vulnerability status observed on the commissioning date because the installed product will encounter future vulnerabilities, component end-of-life events, cryptographic deprecation, dependency changes, and evolving attacker capability throughout its operating life.
Support lifetime is a cybersecurity parameter
Industrial equipment commonly remains operational substantially longer than ordinary enterprise computing platforms, which creates a structural problem when the product’s useful process lifetime exceeds the period during which ordinary supplier security remediation remains available.
Let T_{\mathrm{use}}(p) denote the expected operational lifetime of product p and T_{\mathrm{support}}(p) the period during which ordinary supplier security remediation is available. The unsupported-use interval is
A positive value does not automatically mean that continued operation is impossible or unsafe, but it identifies a lifecycle interval during which newly discovered product vulnerabilities cannot be assumed to receive normal supplier remediation. That interval must be addressed deliberately through extended support, exposure reduction, compensating architecture, migration planning, replacement capability, or explicit residual-risk acceptance.
The EU Cyber Resilience Act makes support duration an explicit part of product cybersecurity and requires manufacturers to consider expected use and other statutory factors when determining support periods; the Regulation also recognizes that products intended for industrial environments, including industrial control systems, may remain operational substantially longer than ordinary digital products.88
Support duration should therefore be evaluated at procurement time as a cybersecurity design parameter rather than discovered shortly before the supplier announces end of support.
SBOMs provide visibility, not proof of security
Software composition is another supply-chain problem because industrial products frequently incorporate operating systems, libraries, runtimes, cryptographic components, embedded applications, and other third-party software whose vulnerabilities may subsequently affect the deployed product.
A Software Bill of Materials can help an operator determine whether a newly disclosed vulnerable component is present in a particular product family or installed version. Without composition information, operators may be unable to establish whether a vulnerability affecting a common library, embedded runtime, operating system, or protocol stack has any relevance to their installed fleet.
CISA maintains SBOM resources intended to improve software-component transparency, while Secure by Demand encourages OT purchasers to seek composition information relevant to software and hardware dependencies.8990
An SBOM is nevertheless composition evidence, not a security verdict. It does not by itself establish whether vulnerable code is reachable, whether the product is securely configured, whether the component is exploitable in the deployed architecture, whether the supplier can remediate it, whether compensating controls constrain resulting authority, or whether the running artifact actually corresponds to the declared component inventory.
The operational value of an SBOM therefore depends on its integration with vulnerability management, product version identification, asset inventory, supplier advisories, configuration management, and architecture-specific exposure analysis.
Supplier due diligence extends beyond product features
The product manufacturer is rarely the complete supply chain. Industrial equipment can depend on contract manufacturers, silicon vendors, embedded operating-system suppliers, open-source projects, certificate authorities, cloud providers, telecommunications companies, engineering subcontractors, maintenance organizations, and lower-tier suppliers that the asset owner may never contract with directly.
NIST SP 1326, finalized in July 2026, structures cybersecurity supply-chain due diligence around dimensions including Foreign Ownership, Control, or Influence; provenance; resilience; foundational cybersecurity practices; and supply-chain tiers.91 These dimensions address a different question from controller hardening or component capability because they examine whether the organizations and dependencies behind the product remain trustworthy, supportable, and resilient. The product-security question is:
What security capabilities does the product provide?
The supplier-risk question is:
What organizations, jurisdictions, services, processes, and lower-tier dependencies must remain trustworthy for that product to remain secure, maintainable, and recoverable?
Both questions are necessary in high-consequence OT. A technically strong controller can still create material operational risk if essential update infrastructure, remote administration, certificate issuance, engineering software, or support services are controlled through dependencies the operator has neither identified nor constrained.
Procurement should request testable evidence
Security properties become materially more useful when they are expressed as acquisition requirements and acceptance evidence rather than as marketing labels. Procurement should therefore translate architecture-level requirements into supplier evidence that can be reviewed before product selection and, where appropriate, validated during factory acceptance, integration testing, commissioning, or lifecycle review.
Contractual support dates, EOL policy, extended-support options, migration path
Unsupported-use interval
SBOM/HBOM or equivalent composition evidence
Machine-readable dependency inventory, versioning method, update process
Hidden component dependencies
Owner autonomy
Local configuration, recovery, maintenance, evidence export, and emergency-operation capabilities
Unnecessary supplier dependency
Remote-support control
Disablement, explicit enablement, local approval, MFA, logging, target restriction, revocation
Permanent supplier conduit
Secure development lifecycle
IEC 62443-4-1 certification or comparable auditable process evidence
Recurrent product-development defects
Component security capability
IEC 62443-4-2 capabilities relevant to the system architecture
Missing native security properties
Secure decommissioning
Credential and key revocation, configuration sanitization, trust-store removal, disposal procedure
Residual trust and information leakage
Table 11: Product-security claims should be translated into procurement evidence and acceptance criteria that expose product limitations before they become architectural dependencies.
The purpose is not to require every product to implement every security mechanism. It is to expose limitations early enough that the system architect can decide whether they are acceptable, whether they can be compensated safely, whether they require additional operational controls, or whether the product should be rejected for the intended use.
The long-term objective is lifecycle governability
The deeper procurement objective is lifecycle governability: the owner should retain enough technical, contractual, and operational authority to understand, constrain, maintain, observe, recover, replace, and eventually retire the product throughout the period in which the physical process depends on it.
For product p, let \operatorname{Governable}(p,t) mean that, at time t, the owner retains sufficient capability to identify deployed state, authenticate and constrain access, observe security-relevant activity, determine supported software and component state, apply or compensate for vulnerability remediation, recover authoritative configuration, replace or migrate the product, and revoke trust during decommissioning.
Governability can be supplied partly by native product capabilities, partly by supplier support, partly by contracts, and partly by the surrounding architecture. What matters is that the capability does not disappear silently while the industrial mission remains dependent on the product.
The strongest supply-chain question is therefore not:
Who manufactured this controller?
It is:
Which product, supplier, software, service, maintenance, update, identity, and remote-access dependencies must remain trustworthy throughout the industrial lifecycle: and does the owner retain enough technical and contractual authority to constrain, observe, recover, replace, and ultimately retire them?
That is the supply-chain counterpart to the capability and dependency models developed throughout this article.
Regulation, governance, and critical-infrastructure accountability
Technology determines which cybersecurity properties can be engineered; governance determines who is accountable for selecting those properties, implementing them, maintaining them, testing them, reviewing them, and producing evidence that they still exist. A technical mechanism becomes an effective control only when it operates inside an organizational system that assigns responsibility and reacts when the evidence no longer supports the intended security property.
For control c, define a simple governance predicate:
The conjunction expresses a genuine assurance condition rather than a maturity score. A policy without implementation is not an operating control, an implementation without evidence provides weak assurance that the required property exists, a control without accountable ownership tends to drift, and a control that is never reassessed can remain formally present long after the architecture, supplier, threat, or physical mission has changed.
Governance is therefore the mechanism that makes cybersecurity properties persistent over time rather than transient features of a commissioning project.
NIS2 makes cybersecurity a management obligation
Directive (EU) 2022/2555 requires covered essential and important entities to implement appropriate and proportionate technical, operational, and organizational cybersecurity risk-management measures.92 Article 21 requires an all-hazards approach and requires the measures to account for factors including risk, exposure, likelihood, severity, implementation cost, the state of the art, and applicable standards.
Its minimum subject areas include risk analysis and information-system security, incident handling, business continuity, backup management, disaster recovery, crisis management, supply-chain security, acquisition and maintenance security, vulnerability handling and disclosure, assessment of risk-management effectiveness, cyber hygiene and training, cryptography, human-resources security, access control, asset management, multifactor or continuous authentication where appropriate, and secured communications where appropriate.93
The all-hazards framing is particularly relevant to OT because cybersecurity cannot be separated completely from physical and environmental dependencies. Loss of power, telecommunications, physical access, environmental control, personnel availability, supplier services, or safety mechanisms can affect the same mission that cyber controls are intended to protect, while cybersecurity measures can themselves alter process availability and safe operation.
Article 20 places accountability explicitly on the management body. Management bodies must approve the Article 21 cybersecurity risk-management measures and oversee their implementation, while Member States must ensure appropriate training for members of those bodies, subject to the Directive’s national implementation and liability framework.94
This does not require directors to understand PLC programming syntax, but it does require governance capable of obtaining credible evidence about questions such as which OT assets are externally reachable, which third parties possess process-level authority, which controllers or operating systems are unsupported, which external dependencies are mission-critical, which sites have tested isolation procedures, whether critical processes can operate safely after loss of remote connectivity, whether recovery baselines have actually been restored during exercises, which safety and ordinary control systems share cyber dependencies, and whether high-consequence engineering actions are observable.
These are governance questions because they concern organizationally accepted exposure and resilience even though their answers depend on engineering evidence.
Reporting obligations depend on detection capability
NIS2 creates a direct operational relationship between governance and observability. For significant incidents, the general staged reporting framework includes an early warning within 24 hours after awareness, an incident notification generally within 72 hours after awareness, and a final report generally no later than one month after the incident notification, subject to the Directive’s rules for continuing incidents and other specific cases.95
The critical concept is awareness. Reporting obligations cannot be operationalized effectively when the organization lacks the technical capability to determine that a significant incident has occurred, establish its scope, identify affected services, or measure material consequence.
The detection architecture developed earlier therefore becomes part of the governance capability. An organization that cannot observe privileged remote sessions, unexpected zone crossings, controller programming, firmware changes, unauthorized identities, or material process impact may also struggle to determine whether the legal significance threshold has been crossed or which facts can be reported with confidence.
Logging, monitoring, controller-state evidence, and incident correlation are consequently not merely SOC conveniences. In regulated critical infrastructure, they can become part of the evidentiary capability through which the organization discharges statutory notification, investigation, and accountability duties.
ENISA guidance operationalizes NIS2, but its direct scope matters
ENISA’s 2025 NIS2 Technical Implementation Guidance provides implementation examples, evidence suggestions, and mappings to other cybersecurity frameworks.96 Its direct context, however, is the technical and methodological requirements established by Commission Implementing Regulation (EU) 2024/2690 for specified categories of digital-infrastructure, ICT service-management, and digital-provider entities.
It should therefore not be represented as though it were a universally binding OT implementation specification for every energy, water, manufacturing, or transport organization falling within NIS2. Its value outside its direct scope is methodological: it illustrates how abstract legal obligations can be translated into implementation expectations, evidence, control mappings, and assurance questions.
The scope distinction matters because governance should distinguish binding legal requirements, authoritative implementing measures, sector-specific obligations, standards, and useful non-binding implementation guidance rather than collapsing them into one undifferentiated compliance checklist.
The EU electricity sector has an additional cybersecurity network code
Electricity has an additional Union-level sectoral cybersecurity instrument. Commission Delegated Regulation (EU) 2024/1366 establishes a network code for cybersecurity aspects of cross-border electricity flows and addresses matters including common minimum cybersecurity requirements, risk assessment, monitoring, reporting, crisis management, and coordination.97
Its significance is architectural as well as regulatory because electricity is a geographically distributed cyber-physical system. A generating station may implement strong local segmentation and controller security while remaining dependent on control centers, transmission-system communications, market and scheduling platforms, telecommunications, neighboring system operators, balancing systems, and external operational coordination.
The security object is therefore not always one plant or one legal entity. Capability and dependency chains can cross organizational and jurisdictional boundaries, and systemic risk can arise from interactions among individually well-controlled organizations. The network code reflects that characteristic by addressing cybersecurity associated with cross-border electricity flows rather than assuming that each operator constitutes an independent security domain.
NIS2 and the Cyber Resilience Act regulate different objects
NIS2 and the Cyber Resilience Act address overlapping cybersecurity risk from different legal directions. NIS2 primarily imposes risk-management, governance, continuity, supply-chain, and reporting obligations on covered entities and services, while the Cyber Resilience Act establishes horizontal cybersecurity requirements for products with digital elements and obligations for manufacturers and other economic operators placing those products on the Union market.98
Their relationship mirrors the architectural division developed earlier. The product supplier must make defensible security capabilities available and sustain them through the product lifecycle; the system integrator or service provider must compose those capabilities into a coherent system; the asset owner and operator must configure, govern, monitor, maintain, and recover the deployed architecture; and the wider supply chain must remain manageable throughout the period in which the industrial mission depends on it.
IEC 62443 spans the same lifecycle through separate parts addressing asset owners, service providers, system design, secure product development, and component capabilities, but the legal frameworks and the standard remain different instruments with different scopes.
A CRA-conforming controller does not make the plant NIS2-compliant, and a NIS2 risk-management program does not create security capabilities that are technically absent from the controller. Product regulation and operator governance are complementary because system security requires both.
CRA applicability dates must be kept distinct
The Cyber Resilience Act uses a staged application schedule. Using the article’s reference date of September 5, 2026, Chapter IV and Articles 35–51 have applied since June 11, 2026; the Article 14 reporting obligations will apply from September 11, 2026; and the Regulation generally applies from December 11, 2027.99
Article 14 is therefore not yet applicable at the article’s stated reference date. From September 11, 2026, manufacturers will have staged notification obligations concerning actively exploited vulnerabilities and severe incidents affecting the security of products with digital elements under the conditions specified by the Regulation.100
For actively exploited vulnerabilities, the staged mechanism includes an early warning within 24 hours and, unless the necessary information has already been provided, a more detailed notification within 72 hours. Severe product-security incidents have an analogous staged reporting structure under Article 14.
This creates a manufacturer-side reporting plane that is distinct from the operator-side incident-reporting obligations under NIS2. One industrial incident can therefore create different duties for the affected operator, the product manufacturer, and potentially other supply-chain or service entities under different triggers, timelines, and reporting channels.
Governance consequently needs to determine not merely whether an incident is reportable, but which legal entity must report which facts, under which regime, after which legal trigger has occurred.
Product support periods become part of regulatory product security
The CRA also reinforces the lifecycle argument developed in the preceding supply-chain section. Manufacturers must determine support periods by considering expected product use and other statutory factors, with a general minimum period of five years unless the product is expected to be used for less than five years.101
The Regulation specifically recognizes that products used in industrial environments, including industrial control systems, can remain operational substantially longer than ordinary digital products. The historical mismatch between industrial operating life and ordinary software-support cycles can therefore no longer be treated exclusively as an asset-owner maintenance problem; manufacturer support, vulnerability handling, and lifecycle planning increasingly form part of regulated product security as well.
For OT procurement, this reinforces the support-gap model introduced earlier. Expected operating life, guaranteed support period, extended-support mechanisms, migration options, and replacement feasibility should be evaluated together before the product becomes embedded in a process whose physical life may extend far beyond ordinary IT refresh cycles.
U.S. electricity regulation combines reliability standards and supply-chain authority
The U.S. Bulk-Power System uses a different legal and institutional architecture. Under section 215 of the Federal Power Act, FERC approves mandatory reliability standards developed through NERC for applicable users, owners, and operators of the Bulk-Power System.102
The NERC Critical Infrastructure Protection standards are therefore fundamentally different from voluntary cybersecurity guidance because they create mandatory reliability obligations for entities and systems within their applicability criteria. In March 2026, FERC approved CIP-003-11, strengthening baseline cybersecurity requirements for low-impact BES Cyber Systems, including remote-user password protections and intrusion-detection requirements, and also approved CIP-002-8 with an updated control-center definition affecting categorization.103
Supply-chain cybersecurity is addressed through standards including CIP-013. As of September 2026, CIP-013-2 is the mandatory version subject to enforcement, while CIP-013-3 is subject to future enforcement with a U.S. effective date of October 1, 2028. The CIP-013 requirements address documented cybersecurity supply-chain risk-management planning covering relevant procurement and vendor risks, including vendor security events, coordination of incident response, termination of vendor access, vulnerability disclosure, software integrity and authenticity, and vendor remote access.104
This maps closely to the architecture developed above because a supplier relationship ceases to be merely commercial when the supplier can deliver executable software, exercise remote authority, distribute updates, manage credentials, or otherwise influence the trustworthiness of BES Cyber Systems.
Executive Order 14420 adds a national-security supply-chain layer
On August 26, 2026, Executive Order 14420 declared a national emergency concerning foreign supply-chain risks to the U.S. bulk-power system.105 The order is technically notable because it defines the relevant security boundary broadly: its transaction provisions can reach physical bulk-power equipment together with associated critical components, software, firmware, digital services, maintenance services, and remote-access capabilities.
Its definition of bulk-power system electric equipment expressly includes classes such as battery energy storage systems, programmable logic controllers, remote terminal units, intelligent electronic devices, distributed control systems, and safety instrumented systems. The order also gives the Secretary of Energy authority, under the stated conditions, to impose requirements concerning existing foreign-manufactured or operated equipment, including measures to identify, isolate, monitor, secure, disconnect, replace, or remove it.106
An equally important OT constraint appears in the requirement to consider reliability, safety, availability of secure replacements, and continuity of essential service before directing isolation, disconnection, replacement, or removal.107
That condition mirrors the cyber-physical admissibility principle developed earlier. Even where equipment presents a cybersecurity or national-security concern, mitigation cannot be evaluated independently from the consequences of removing the equipment from a physical mission. Cybersecurity intervention must remain constrained by safe and reliable process operation.
The UK is moving toward broader downstream-energy baselines
On August 5, 2026, the UK Department for Energy Security and Net Zero and Ofgem published the government response to their consultation on cybersecurity regulation in downstream gas and electricity.108 The government stated its intention to review applicability of the Network and Information Systems Regulations 2018 in the downstream gas and electricity sector and to develop baseline cyber-resilience requirements for all Ofgem licensees.
The legal status matters. These statements describe announced policy direction and intended next steps; they should not be represented as though every proposed baseline requirement were already legally binding on every Ofgem licensee.
The distinction between current obligation and declared regulatory direction is essential in comparative analysis because emerging policy can be highly relevant to architecture, procurement, and investment planning without yet constituting an enforceable requirement.
Australia uses an all-hazards critical-infrastructure model
Australia’s Security of Critical Infrastructure Act 2018 establishes a broader critical-infrastructure risk framework.109 For applicable responsible entities, the framework includes critical-infrastructure risk-management programs, cyber-incident notification requirements, and enhanced cybersecurity obligations for specified systems of national significance.
The risk-management-program model is explicitly broader than cyberattack alone because relevant entities must identify and address material hazards capable of affecting critical-infrastructure assets. Cybersecurity consequently sits inside a wider resilience model incorporating physical security, personnel, supply chain, operations, environmental conditions, and other dependencies.
That structure is particularly compatible with OT because cyber, physical, personnel, supplier, and environmental failures can converge on the same process state and the same physical mission even when they originate in different organizational disciplines.
Different regimes increasingly converge on similar governance properties
Table 12 summarizes the principal regulatory structures discussed here.
Products with digital elements and relevant economic operators
Secure-product requirements, vulnerability handling, support periods, conformity framework, manufacturer reporting
EU electricity cybersecurity network code
Cybersecurity aspects of cross-border electricity flows
Sector-specific risk assessment, minimum requirements, monitoring, reporting, and crisis coordination
FERC/NERC CIP
Applicable Bulk Electric System entities and cyber systems
Mandatory cybersecurity reliability standards
U.S. Executive Order 14420
Foreign supply-chain risk involving bulk-power system electric equipment
Transaction restrictions, monitoring, conditions on use, isolation, disconnection, replacement, and supply-chain controls
UK 2026 downstream-energy reform direction
Downstream gas and electricity regulatory scope
Planned review of NIS applicability and proposed baseline cyber-resilience requirements for Ofgem licensees
Australia SOCI
Critical-infrastructure assets and responsible entities
All-hazards risk-management programs, incident reporting, governance, and enhanced obligations for designated systems
Table 12: Different legal regimes use different mechanisms but increasingly converge on accountable risk management, resilience, supply-chain governance, evidence, and lifecycle responsibility.
The legal architectures remain materially different and should not be collapsed into a claim that these jurisdictions have adopted one common cybersecurity model. Their regulatory objects, enforcement structures, applicability criteria, reporting mechanisms, and relationships with sector-specific law remain distinct.
The convergence occurs instead at the level of governance properties: risk must be identified, responsibility assigned, preventive and resilience measures implemented, suppliers and dependencies governed, incidents detected and reported, continuity and recovery demonstrated, and evidence retained showing whether the controls actually operate.
Compliance is not equivalent to resilience
The central governance danger is replacing engineering with compliance. Conformity with a regulatory requirement, contractual clause, framework control, or audit procedure can provide valuable discipline and accountability, but documentary compliance is not logically equivalent to cyber-physical resilience.
A firewall appearing in an audit inventory does not establish that compromise of its management plane cannot collapse several zones; an approved backup policy does not demonstrate that the turbine-control environment can be reconstructed from trusted state; a vendor-access procedure does not prove that dormant supplier accounts have actually been removed; an IEC 62443-certified component does not establish the achieved security properties of the integrated plant; and a NIS2 policy does not prove that the organization can detect an unauthorized PLC mode change.
Compliance provides value when it creates disciplined requirements, ownership, traceability, testing, and evidence. It becomes dangerous when documentary satisfaction substitutes for determining whether the intended cyber-physical property exists in the deployed architecture.
The assurance relationship is therefore better represented by the traceability structure in Figure 11 rather than as an implication equation.
%%{init: {"theme": "neo", "look": "handDrawn", "layout": "elk"}}%%
flowchart TD
A[Legal or risk<br/>obligation]
B[Security<br/>objective]
C[Engineering<br/>requirement]
D[Implemented<br/>control]
E[Operational<br/>evidence]
F[Review and<br/>adaptation]
A -->|translated into| B
B -->|made testable as| C
C -->|realized by| D
D -->|produces| E
E -->|supports assurance and challenge| F
F -->|updates objectives and assumptions| B
Figure 11: Governance assurance chain from legal or risk obligation to observable evidence and reassessment. Rectangles represent governance or engineering artifacts and ordinary arrows represent traceability and dependency relationships, not attacker-capability escalation.
Every relationship in that chain can fail. A legal requirement can be translated into the wrong security objective, a sound objective can produce an inadequate engineering requirement, a correct requirement can be implemented incorrectly, a technically valid implementation can drift, and evidence can be collected indefinitely without anyone determining whether it still demonstrates the intended property.
Governance therefore requires traceability from obligation to observable system behavior and back again through review.
Governance evidence should correspond to engineering properties
The formal models developed throughout this article illustrate what stronger governance evidence can look like because they describe properties of the deployed architecture rather than the mere existence of documentation.
For segmentation, the relevant evidence is not only that firewall rules exist but whether the permitted inter-zone flow set corresponds to the mission-required flow set defined earlier. For remote access, the evidence is not that MFA or a ZTNA product has been purchased but whether remote principals can reach only the resources and exercise only the operations required by their approved role.
For containment, evidence is not the existence of an isolation procedure but whether a tested response can make the relevant high-consequence derivation infeasible while preserving safe process control. For recovery, evidence is not the presence of backup files but whether the required recovery artifacts are valid, complete, sufficiently current, trustworthy, and demonstrably restorable within the operational constraints of the mission.
For controller integrity, evidence is not merely that engineering project files are archived but whether the deployed controller state can be shown, using product-appropriate verification, to correspond to the independently approved state. For resilience, evidence is not aggregate infrastructure uptime but whether credible degraded cyber-operational states still preserve mission viability or allow the process to reach an approved safe state.
This is how governance reconnects to engineering. The regulatory question
Are appropriate cybersecurity measures in place?
must eventually be answered with evidence about the actual industrial architecture, its trusted states, its degraded modes, and the behavior of the physical mission.
Accountability closes the architectural loop
The capability-inference model introduced earlier asks whether a plausible initial foothold can derive a high-consequence capability or state. Engineering seeks to make those derivations infeasible; detection determines whether the supposedly constrained relationships, actions, and states are behaving as intended; recovery ensures that compromise does not permanently destroy trustworthy mission capability; supply-chain governance determines whether the necessary product and supplier properties remain available throughout the lifecycle.
Regulation and organizational governance add the final element: someone must be accountable for ensuring that each required property continues to exist, for reviewing the evidence that supports it, and for acting when that evidence fails. The final governance question is therefore not:
Which cybersecurity frameworks and regulations apply to this plant?
It is:
For every high-consequence derivation, who owns the requirement intended to interrupt it, which technical or operational control implements that requirement, what evidence demonstrates that the control still provides the intended property, and who is accountable for acting when that evidence no longer supports the claim?
That is the point at which compliance becomes engineering assurance rather than documentation.
A reference defense-in-depth architecture for modern OT
The preceding analysis can now be assembled into a single architectural proposition:
A plausible compromise of one ordinary entry point should not, by itself, compose into high-consequence process authority.
This is more precise than requiring every individual component to resist compromise indefinitely. Modern OT environments contain externally reachable identities, remote-support relationships, engineering systems, supervisory platforms, gateways, distributed sites, and supplier dependencies; some of those elements will eventually be compromised. The architectural question is whether one such compromise is sufficient to derive the authority required to manipulate a high-consequence controller state or produce an unacceptable physical effect.
Let \mathcal{E}_{\mathrm{entry}}\subseteq V denote the subset of attacker capabilities corresponding to plausible compromise of ordinary entry points, such as one remote identity, enterprise endpoint, vendor relationship, remote-access session, gateway, remote site, or ordinary supervisory or management system. The set deliberately excludes initial conditions that already contain the high-consequence capability being protected against: if the assumed starting condition is unrestricted compromise of the safety controller with complete process authority, the architecture cannot meaningfully claim to prevent that same starting condition from providing high-consequence authority.
For architectural state z, deployed controls K, and high-consequence capability set H, the desired single-foothold containment property is
C_{z,K}^{*}\left(\{v\}\right)
\cap
H
=
\varnothing
\qquad
\forall v\in\mathcal{E}_{\mathrm{entry}}.
\tag{53}
The objective is therefore not:
Nothing can ever be compromised.
It is:
One ordinary compromise should not be sufficient to derive uncontrolled process authority or consequence.
Achieving that property requires several independently governed conditions between an external or ordinary foothold and high-consequence control: identity must be attributable, the originating endpoint must be sufficiently trusted, privileged admission must be mediated, only the required conduit and target should become reachable, the resulting session must expose only the required engineering or operational authority, controller-native authorization must further bound what can be changed, and high-consequence controller state must remain protected by controls appropriate to its physical effect. These are security gates, not necessarily separate appliances. Their value comes from preventing compromise of one condition from automatically satisfying all the others.
Authority domains are more useful than arbitrary network levels
Traditional automation hierarchies remain useful for describing industrial function and data flow, but numbered levels become misleading when they are treated as though they automatically constitute security boundaries. Two systems occupying the same functional automation level can have radically different administrative privileges, process consequences, supplier dependencies, recovery requirements, or safety significance, while systems on different levels may legitimately share narrowly constrained relationships.
The more useful architectural question is therefore:
Which assets share comparable authority, trust assumptions, consequences, and recovery dependencies?
A representative OT environment can distinguish conceptual domains for external suppliers and remote users, enterprise systems, the industrial DMZ, OT identity and security management, privileged engineering, supervisory operations, ordinary process cells, high-consequence control, independent safety or protection, distributed remote sites, monitoring and evidence, and recovery trust.
These conceptual domains need not map one-to-one to VLANs, physical servers, or network appliances. A small facility may implement several functions using relatively few systems, while a utility may divide one conceptual authority domain into dozens of security zones. What matters is that systems whose compromise produces materially different authority or consequence are not treated as one undifferentiated OT network.
This is consistent with the IEC 62443 principle developed earlier: zones should be derived from security requirements, risk, trust, function, and consequence rather than mechanically inherited from an automation hierarchy.
Crossing functions should remain distinct at the IT/OT boundary
Several fundamentally different functions legitimately cross an enterprise/OT boundary, and they should not be collapsed into one generic routed relationship. At minimum, the architecture should distinguish privileged administration, operational-data export, controlled file or content import, and security-evidence export, because each moves a different object across the boundary and creates a different trust requirement.
Privileged administration carries operational or engineering authority toward the process. Historian or telemetry replication carries process information away from it. Controlled file import introduces content that may later become executable, configuration-changing, or otherwise authoritative. Security-evidence export moves observations toward an analytical environment or SOC but ordinarily should not create corresponding administrative authority in the reverse direction.
Historian export therefore does not require enterprise reachability to controllers, and security-log export does not require the SOC to possess administrative authority over PLCs. Likewise, the existence of a legitimate software-import path does not justify general bidirectional file sharing between enterprise and engineering workstations.
The industrial DMZ should mediate these crossing functions separately so that each conduit carries only the direction, protocol, content, identity, and authority required by its mission. A single broad VPN or generic firewall opening that combines administration, file transfer, telemetry, and evidence export destroys precisely the distinctions the boundary is intended to enforce.
Supervisory authority and engineering authority should remain distinct
Operator authority and engineering authority are generally different kinds of industrial authority rather than nested privilege levels. An operator may legitimately start or stop equipment, acknowledge alarms, or change a production setpoint while having no reason to modify controller logic; an engineer may legitimately download an approved project or alter controller configuration while having no operational need to start the process.
For operator principal p and engineering principal q, let \mathcal{O}_{\mathrm{HMI}}(p,z) and \mathcal{O}_{\mathrm{ENG}}(q,z) denote the industrial operations available through the supervisory and engineering environments respectively. Their least-authority requirements are role-specific:
The two sets may overlap where both roles legitimately require the same operation, but neither should be inferred from the other. Ordinary supervisory access should not automatically permit firmware replacement, modification of protection settings, creation of privileged controller identities, unrestricted project download, or modification of safety logic; possession of an engineering identity should similarly not create unrestricted process-operating authority.
Separating these authority domains limits the consequence of compromise because a compromised HMI does not automatically become a programming station, while compromise of an engineering account does not necessarily provide every ordinary production command.
Process cells need explicit principal–operation relationships
For process cell Z_i, let \mathcal{P}_i denote the principals or systems that can legitimately interact with the cell and let \mathcal{O}_i denote the industrial operations relevant to that cell. Their Cartesian product \mathcal{P}_i\times\mathcal{O}_i is not an authorization model because it would implicitly pair every principal with every operation.
Instead, define the mission-required authorization relation \Pi_i^{\mathrm{mission}}\subseteq\mathcal{P}_i\times\mathcal{O}_i. The deployed architecture should permit exactly that relation:
A historian may therefore be authorized for a defined set of reads, an HMI for a bounded set of process operations, and a privileged engineering workstation for explicitly governed engineering functions. The fact that all three legitimately communicate with the same controller does not mean that they should possess equivalent controller authority.
Where a firewall can constrain communicating peers but cannot interpret industrial operation semantics, the remaining authorization requirement must be enforced elsewhere through controller-native authorization, protocol-aware gateways, dedicated engineering hosts, application roles, programming windows, or stronger zone separation. This is the practical composition of IEC 62443 FR 2, Use Control, and FR 5, Restricted Data Flow: the conduit determines which relationship can exist, while downstream controls determine which authority that admitted relationship can exercise.
Higher consequence should imply stronger path restriction
It is not mathematically valid to require the reachable-principal set of every high-consequence zone to be a strict subset of the corresponding set for every lower-consequence zone, because zones perform different functions and can require legitimate peer sets that are not directly comparable. The architectural principle is normative rather than ordinal:
Higher consequence should require stronger justification for every reachable principal, conduit, protocol, and industrial operation.
For each zone Z_i, the excess authorization relation is \Delta\Pi_i=\Pi_i^{\mathrm{allowed}}\setminus\Pi_i^{\mathrm{mission}}. Least authority requires \Delta\Pi_i=\varnothing regardless of consequence level; what changes for high-consequence zones is the required strength, independence, and evidentiary quality of the controls that enforce the permitted relation.
High-consequence control can therefore justify fewer administrative principals, dedicated engineering endpoints, explicit programming windows, independently governed privileged identities, stronger controller-native protection, narrower protocol capabilities, hardware-enforced or physically independent boundaries where warranted, and more stringent monitoring of changes in controller authority and state.
The objective is not to accumulate controls for their own sake. It is to ensure that increasing physical consequence is accompanied by greater independence among the conditions that must all hold before high-consequence authority becomes available.
Safety independence is a dependency property
Where independent safety or protection functions exist, cybersecurity architecture should preserve the independence on which the safety design relies. The earlier model distinguished the cyber dependencies of ordinary control, D_C, from those of safety or protection, D_S, and identified D_{\mathrm{shared}}=D_C\cap D_S as the set of candidate common-mode dependencies requiring explicit examination.
The architectural objective is not necessarily to make D_{\mathrm{shared}} empty under every design, but to minimize unnecessary shared dependencies and determine what authority compromise of each remaining common element can derive. This is stronger than assigning the BPCS and SIS different IP addresses because network separation can coexist with common-mode failure through the same engineering workstation, privileged identity, directory service, remote-access broker, switch or firewall management plane, firmware repository, supplier maintenance mechanism, administrative personnel, or recovery infrastructure.
A safety controller can therefore be topologically separate while remaining cyber-dependent on the same authority plane as ordinary control.
The relevant architectural question is:
Which plausible single cyber compromise could simultaneously undermine ordinary control and the independent protection intended to constrain its consequence?
Where such a compromise exists, the safety boundary must be evaluated as a common-mode dependency rather than assumed to be independent because the safety controller occupies another network segment.
Remote sites should not inherit trust from shared transport
Distributed OT frequently uses private APNs, MPLS, carrier VPNs, SD-WAN, leased telecommunications, or other shared transport. These services can provide useful communications isolation from the public Internet, but carrier membership is not an endpoint identity and should not automatically create lateral trust among every connected industrial site.
For remote sites S_i and S_j, when no mission relationship exists between them, the least-functionality principle developed earlier requires \mathcal{F}_{ij}^{\mathrm{mission}}=\varnothing and therefore \mathcal{F}_{ij}^{\mathrm{allowed}}=\varnothing. Shared transport should provide connectivity to explicitly authorized central or peer services, not transform the carrier network into one flat OT security zone.
A remote site should therefore ordinarily communicate through a controlled site gateway toward defined central systems, with site-to-site relationships enabled only when the mission explicitly requires them. This prevents compromise of one peripheral site from becoming a routing opportunity toward every other facility simply because all of them share the same telecommunications service.
Management infrastructure is itself high-consequence OT
Infrastructure that establishes or administers security boundaries can have authority greater than many of the controllers it protects. Firewalls, routers, managed switches, hypervisors, privileged-access brokers, PKI services, OT identity systems, ZTNA connectors, virtualization platforms, software-deployment services, and network-management systems can redefine trust relationships for entire portions of the architecture.
Compromise of one such platform can therefore collapse several apparently independent boundaries simultaneously. A firewall provides segmentation only while its management plane remains trustworthy; an identity system enforces separation only while privileged role assignment and credential issuance remain trustworthy; a virtualization platform can undermine multiple isolated servers if one administrative plane controls them all.
Management infrastructure should consequently be treated as a high-authority OT domain and protected accordingly through dedicated privileged identities, PAWs, restricted management conduits, independently controlled or out-of-band access where justified, protected configuration backups, privileged-action logging, strong monitoring, controlled software and firmware lifecycles, and separation from ordinary enterprise administration.
This directly addresses the shared-infrastructure failure identified in the taxonomy: logical separation implemented by one common management authority can remain vulnerable to common-mode administrative compromise.
Recovery is a separate trust domain
Recovery material should not be treated as another ordinary production resource. The recovery architecture developed earlier requires production compromise not to imply loss of the trusted baselines, credentials, tools, and configuration material on which reconstitution depends.
That property requires partial independence in administrative identities, credentials, repositories, backup infrastructure, privileged-access mechanisms, destructive-operation authorization, and management paths. The exact degree of separation depends on consequence and operational scale, but the purpose of the recovery domain is fundamentally different from daily production: it must remain trustworthy precisely when ordinary production trust can no longer be assumed.
Recovery infrastructure should therefore be designed as a distinct trust domain rather than as additional storage attached to the same administrative environment that operates production.
A reference architecture organized around constrained authority
Figure 12 integrates these principles into a conceptual architecture. It is deliberately not a mandatory topology and not a Purdue diagram: it represents authority domains, mediation mechanisms, privileged relationships, high-consequence control, monitoring, and recovery trust.
%%{init: {"theme": "neo", "look": "handDrawn", "layout": "elk"}}%%
flowchart TB
subgraph EXT[External / supplier domain]
V[Vendor / integrator]
R[Remote engineer]
TE[Trusted administrative<br/>endpoint / PAW]
CL[Cloud / managed service]
CAR[Carrier network]
end
subgraph ENT[Enterprise domain]
EU[Enterprise users]
EID[Enterprise identity]
BA[Business analytics]
SOC[Enterprise / national SOC]
end
AUTH{Named identity, MFA,<br/>endpoint trust and contextual<br/>admission conditions satisfied}
ZT[/ZTNA / remote-access<br/>policy mediation/]
V -->|authorized supplier activity| TE
R -->|privileged activity originates here| TE
TE -->|authenticated administrative request| ZT
EID -->|identity assertions| ZT
AUTH -.->|required admission conditions| ZT
subgraph DMZ[Industrial DMZ]
AB[/Privileged access<br/>broker/]
FT[/Controlled file<br/>staging/]
DH[Operational-data<br/>replica]
SG[/Security-evidence<br/>export gateway/]
end
ZT -->|resource-specific admission| AB
CL -->|controlled content import| FT
subgraph SEC[OT identity / security management]
OID[OT privileged<br/>identity service]
PKI[OT PKI / trust<br/>services]
NM[Network / security<br/>management]
end
OID -->|OT privileged identity| AB
PKI -->|trust services| OID
subgraph ENG[Privileged engineering]
PES[[Bounded privileged<br/>engineering session]]
EW[Engineering<br/>workstation]
REP[Approved project /<br/>firmware repository]
end
AB ==>|policy grants bounded privilege| PES
PES -->|session restricted to<br/>approved engineering resource| EW
FT -->|validated content transfer| REP
subgraph SUP[Supervisory operations]
HMI[Operator HMI]
SCADA[SCADA /<br/>control server]
HIST[OT historian]
ALM[Alarm services]
end
HMI -->|operator interaction| SCADA
subgraph C1[Process cell A]
PLC1[PLC]
IO1[Drives / remote I/O]
end
subgraph C2[Process cell B]
PLC2[PLC]
IO2[Drives / remote I/O]
end
subgraph HC[High-consequence control]
PLC3[Critical controller]
end
subgraph SAFE[Independent safety / protection]
SEW[Dedicated safety<br/>engineering]
SIS[SIS / protective<br/>controller]
end
SCADA -->|Constrained operational control| PLC1
SCADA -->|Constrained operational control| PLC2
SCADA -->|Restricted operational control| PLC3
PLC1 -->|Local control| IO1
PLC2 -->|Local control| IO2
EW -->|Time-bounded engineering| PLC1
EW -->|Time-bounded engineering| PLC2
EW -->|Restricted engineering| PLC3
SEW -->|Dedicated safety engineering| SIS
PLC1 -->|Process data| HIST
PLC2 -->|Process data| HIST
PLC3 -->|Process data| HIST
SIS -->|Status / alarms only| ALM
HIST -->|Constrained operational-data export| DH
DH -->|Approved external consumption| BA
subgraph REM[Distributed remote sites]
S1[Remote PLC / RTU A]
G1[/Site gateway A/]
S2[Remote PLC / RTU B]
G2[/Site gateway B/]
CG[/Central OT<br/>WAN gateway/]
end
S1 -->|Site communication| G1
S2 -->|Site communication| G2
G1 -->|Authorized carrier path| CAR
G2 -->|Authorized carrier path| CAR
CAR -->|Shared transport only| CG
CG -->|Defined central OT relationship| SCADA
subgraph MON[Monitoring / evidence]
PS[/Passive protocol<br/>observation/]
EC[OT evidence<br/>collector]
end
AB -.->|access events| EC
EW -.->|engineering events| EC
SCADA -.->|traffic copy| PS
PLC1 -.->|traffic copy| PS
PLC2 -.->|traffic copy| PS
PLC3 -.->|traffic copy| PS
CG -.->|traffic copy| PS
PS -.->|decoded network evidence| EC
EC -->|Constrained evidence export| SG
SG -->|Security evidence| SOC
subgraph REC[Recovery trust domain]
BK[Immutable / offline<br/>backups]
GC[Approved controller<br/>baselines]
NC[Network / security<br/>configurations]
RI[Recovery tools and<br/>credential material]
end
REP -->|Approved baseline replication| GC
NM -->|Protected configuration export| NC
HIST -->|Protected backup| BK
Figure 12: Reference defense-in-depth OT architecture organized around constrained operational authority. Rectangles represent ordinary principals, systems, resources, or repositories; slanted nodes represent mediation, brokering, staging, gateway, or observational mechanisms; braces represent security conditions governing admission; double brackets represent privileged operational or engineering authority. High-consequence control and independent safety are identified by their domain boundaries rather than by special node shapes. Dotted arrows carry observational evidence or contributory security conditions; ordinary arrows represent permitted communication, dependency, or data transfer; thick arrows represent the explicit grant of privileged authority. These arrows describe the reference architecture and are not attacker-capability transitions.
The figure represents authority and trust relationships, not a required count of VLANs, switches, firewalls, or servers. The implementation can vary substantially with process scale and consequence, but the architectural distinctions should remain: privileged engineering is not ordinary supervision, data export is not administrative reachability, safety is not merely another PLC network, carrier connectivity does not create lateral site trust, evidence collection should not become a process prerequisite, and trusted recovery should not share unrestricted production authority.
The most important controls are often absent paths
Security architecture is defined as much by relationships that do not exist as by those that are implemented. A well-designed plant should ordinarily have no direct Internet-to-PLC relationship, no ordinary enterprise-workstation path to PLC programming, no unmanaged vendor-laptop relationship with an OT subnet, and no remote-site-to-remote-site connectivity merely because both sites use the same carrier network.
These are stronger statements than saying that traffic is inspected or that suspicious activity is monitored. If a relationship has no mission requirement, the preferred architecture is normally to make it unavailable rather than to permit it and depend on downstream detection.
In the capability-inference model, an absent conduit, unavailable privilege, disabled protocol function, missing credential relationship, or independently enforced trust boundary can make a rule infeasible and thereby prevent downstream capabilities from entering the closure. The defensive value is therefore structural: the attacker is not merely made less likely to progress; one or more required derivations cease to be available under the modeled architecture. This is why unnecessary reachability should be removed rather than merely observed.
Authority should contract as a session approaches the process
A remote-access path often begins with relatively broad organizational eligibility but should progressively resolve into narrower technical authority. A supplier relationship may identify an organization, authentication resolves that relationship into a named person, endpoint policy restricts the person to an approved device, an access broker creates one bounded session, that session reaches one engineering environment, the engineering environment reaches selected controllers, controller policy permits only defined operations, and change governance restricts those operations to an approved intervention.
To model this, let \mathcal{A}_k(s) denote the set of downstream industrial actions still available to privileged session s after restrictive security gate k, with every \mathcal{A}_k(s) defined over the same universe of possible actions. Where each gate is intended only to restrict the session, the architecture should satisfy \mathcal{A}_{k+1}(s)\subseteq\mathcal{A}_k(s), and therefore
The model is intentionally restricted to gates whose function is to narrow an already admitted session. It should not be generalized to unrelated architectural controls that can introduce new services or dependencies.
Moving toward the physical process should therefore resolve and contract authority, not suddenly expand it. This is the architectural opposite of a flat routed VPN in which successful authentication to one access service can expose an entire industrial subnet and leave protocol or controller semantics as the only remaining constraint.
Defense in depth requires partially independent failure modes
Multiple cybersecurity products do not automatically constitute defense in depth. A remote-access architecture can contain MFA, ZTNA, a firewall, a jump host, and controller authentication while all five ultimately depend on the same enterprise identity service, privileged administrator, credential store, hypervisor, management workstation, or management plane. One common-mode compromise can then invalidate several nominally separate layers.
Defense in depth therefore requires partial independence of failure modes, not merely multiple product names. Where consequence justifies it, architecture can separate enterprise and OT privileged identities, ordinary and safety engineering, production and recovery credentials, production and monitoring administration, remote-access and local emergency-control paths, primary and recovery repositories, or software-enforced and independently enforced boundaries.
Perfect independence is rarely achievable and is not the objective. The stronger formal requirement was established earlier through defensive-loss scenarios: for the modeled losses L\in\mathfrak{L}, the remaining controls should still prevent derivation of high-consequence capability,
The architectural meaning is that loss of one modeled security layer should not satisfy every remaining prerequisite required to reach H. Defense in depth becomes meaningful when the failures that defeat different controls are not all the same failure expressed through different products.
IEC 62443 requirements are realized through interacting mechanisms
Table 13 summarizes how the reference architecture realizes the seven IEC 62443 Foundational Requirements. The mapping is many-to-many: each architectural mechanism can support several FRs, and each FR normally requires several mechanisms distributed across products, system design, operations, and lifecycle governance.
Architectural mechanism
Principal IEC 62443 FR
Architectural effect
Named identities, MFA, endpoint authentication, dedicated OT identities
FR 1 — Identification and Authentication Control
Reduces impersonation, anonymous administration, and ambiguous attribution
Role separation, PAWs, programming windows, operation-specific authorization
FR 2 — Use Control
Restricts what authenticated principals may actually do
Makes misuse, anomalous relationships, and unauthorized state changes observable
Local autonomy, isolation capability, trusted backups, recovery domains, safe reconstitution
FR 7 — Resource Availability
Preserves or restores trustworthy mission capability
Table 13: IEC 62443 Foundational Requirements are realized through interacting architectural mechanisms rather than isolated product controls.
The table also demonstrates why no individual FR constitutes an architecture. FR 5 can prevent an enterprise endpoint from reaching a controller, but where an engineering relationship is legitimately required, FR 1, FR 2, and FR 3 still determine whether that permitted relationship can become uncontrolled controller authority. Likewise, FR 7 should not be reduced to keeping computers online: in the mission-viability model developed earlier, availability concerns preservation or restoration of enough trustworthy capability to maintain the physical mission or move the process to an approved safe state.
A network diagram is only one layer of the architecture
A conventional network diagram captures an important but incomplete representation of OT security because many consequential dependencies do not correspond to routed network edges. Identity systems can authorize sessions, suppliers can control update channels, privileged administrators can govern several supposedly separate zones, recovery repositories can depend on production identities, and monitoring platforms can derive all of their evidence from one compromised source.
where \mathcal{X} is the set of security-relevant architectural elements; E_N represents network reachability; E_I identity and trust dependencies; E_A administrative and operational authority relationships; E_D operational dependencies required by the physical mission; E_M monitoring and evidence dependencies; E_R recovery dependencies; and E_S supplier and external-service dependencies.
Using \mathcal{X} rather than the capability set V is intentional. This architectural graph describes what systems and dependencies exist; the capability-inference model developed earlier describes what attacker capabilities become derivable given those dependencies, an initial foothold, and the architectural-feasibility predicates. They are related models, but they are not the same mathematical object.
A conventional network drawing primarily exposes E_N. Yet an incident can depend simultaneously on a supplier relationship in E_S, an identity dependency in E_I, a reachable gateway in E_N, and an authority relationship in E_A before a controller or mission dependency in E_D is affected. Recovery can similarly fail because production and recovery share identity or administrative dependencies, while monitoring can fail because every evidence source ultimately depends on one compromised controller or management platform. The architectural model therefore asks:
Which reachability, identity, authority, mission, monitoring, recovery, and supplier dependencies exist, and where are common-mode dependencies concentrated?
The capability-inference model asks:
Given that architecture, a plausible initial capability, and the current architectural state, which additional capabilities and high-consequence states are derivable?
The first describes the dependency structure. The second evaluates the security consequence of compromise within that structure.
The reference architecture is an authority-containment architecture
The phrase defense in depth can create the impression that OT security consists primarily of placing several defensive products between the Internet and a PLC. That interpretation is too narrow for modern industrial environments, which may legitimately require cloud services, remote support, enterprise analytics, distributed sites, centralized identity, supplier maintenance, external dispatch, and remote engineering.
The architectural objective cannot therefore be universal disconnection. It is controlled composition of authority.
Each permitted relationship should answer five questions:
Who or what is the principal?
From which sufficiently trusted endpoint may it act?
Which specific resource may it reach?
Which industrial operations may it perform?
Which independent controls prevent that legitimate authority from composing into an unacceptable physical consequence?
For process cell Z_i, the target architecture can be summarized by three complementary properties:
The first property asks whether operational authority is restricted to the principal–operation relationships required by the mission. The second asks whether the modeled compromise scenarios remain separated from high-consequence capabilities and states. The third asks whether the surviving trusted architecture still permits the physical mission to remain viable under the relevant cyber-operational state.
These correspond to the three central engineering questions developed throughout the article:
Is authority limited to what the industrial mission actually requires?
Can a plausible compromise derive unacceptable process authority or consequence?
Can the physical mission remain safe and viable when parts of the digital architecture fail or become untrustworthy?
A reference OT architecture is therefore not best understood as a prescribed arrangement of firewalls, DMZs, VLANs, jump hosts, or Purdue levels. It is an authority-containment architecture whose purpose is to preserve legitimate industrial connectivity while ensuring that high-consequence authority remains deliberately narrow, context-bound, observable, independently constrained, and recoverable.
Limits, trade-offs, and open technical questions
Defense in depth reduces the number of short derivations through which an ordinary digital compromise can accumulate into high-consequence industrial authority, but it does not eliminate cyber-physical risk. The architecture developed in the preceding sections remains exposed to common-mode failure, operational complexity, brownfield constraints, incomplete detection, imperfect assurance, supplier dependence, cryptographic lifecycle problems, and residual attacker capability.
These limitations matter because an architecture can be logically coherent at design time while depending on assumptions that become difficult to preserve over a twenty-year industrial lifecycle. Identity systems change, vendors disappear, engineering tools become obsolete, cryptographic mechanisms age, production configurations drift, operators create exceptions, and apparently independent security controls may turn out to share the same administrative root of trust. Defense in depth should therefore be treated as an architectural method for bounding and decomposing risk, not as a claim that enough layers can make cyber-physical compromise impossible.
Layers can fail together when they share roots of trust
Several security barriers do not constitute independent defense in depth merely because they are implemented by different products. Let D(B_i) denote the set of security-relevant dependencies on which barrier B_i relies. For barriers B_i and B_j, their shared dependency set is
Membership in D_{ij}^{\mathrm{shared}} does not by itself prove common-mode failure. The relevant question is whether a shared dependency possesses sufficient authority over both barriers that its compromise can invalidate their intended separation. A common identity service, privileged administrator, hypervisor, management workstation, certificate authority, configuration platform, or supplier-management path can therefore matter much more than the number of nominally distinct security products deployed.
Examples include several firewalls administered through one compromised privileged identity, PAWs and access brokers joined to the same compromised administrative domain, production and recovery repositories controlled through the same authority plane, ordinary control and safety engineering performed from the same workstation, or multiple monitoring products deriving their evidence from the same compromised controller.
Defense in depth is consequently also a problem of failure-domain diversity. Adding another barrier provides limited assurance when the new layer inherits the dominant root of trust of the layers already present. The meaningful architectural question is not how many controls exist, but how many materially different failures are required to defeat them together.
Security complexity can become insecurity
Cybersecurity mechanisms impose a lifecycle burden that extends far beyond acquisition cost. Every new firewall, privileged-access workflow, certificate hierarchy, segmentation rule, broker, monitoring platform, identity domain, recovery mechanism, and approval process creates additional integration, operational, maintenance, troubleshooting, documentation, training, and governance requirements. Those interactions can themselves become a source of security failure when the architecture becomes too difficult to operate correctly under normal production pressure.
The practical warning signs are familiar: shared emergency credentials appear because named access is too cumbersome; vendor VPNs remain permanently enabled because repeated authorization is operationally inconvenient; firewall exceptions become undocumented because troubleshooting procedures are too slow; USB transfer bypasses the approved staging mechanism; engineers use unmanaged laptops because dedicated workstations are unavailable; approval workflows are routinely bypassed; or supposedly temporary privileged sessions remain active indefinitely.
A formally restrictive architecture can therefore become weaker in practice when legitimate engineering work is so difficult that operators must routinely circumvent the controls to keep the plant running. The objective is not maximum security-control complexity but the simplest architecture capable of enforcing the required security invariants with acceptable operational reliability.
This is an engineering trade-off rather than an argument for weaker controls. Complexity should be removed where it does not create independent security value, while controls protecting genuinely different authority boundaries should remain explicit even when doing so increases operational effort.
Segmentation has latency, availability, and maintenance costs
Segmentation is not free. Firewalls, proxies, industrial gateways, protocol inspection, NAT, authentication mechanisms, and mediation layers can introduce latency, jitter, compatibility problems, additional failure modes, troubleshooting difficulty, and maintenance dependencies. They therefore remain subject to the cyber-physical admissibility principles developed earlier: a control that reduces attacker capability while destabilizing the physical process is not an acceptable OT control.
More segmentation also does not automatically produce greater security. With n zones there can be as many as n(n-1) directed inter-zone relationships; if every zone remains permitted to initiate broad communication with every other zone, increasing the number of VLANs or firewall objects can increase administrative complexity without materially constraining authority.
The meaningful property remains whether the deployed principal–operation and zone-to-zone relationships correspond to mission requirements without unnecessary permissions. A small number of rigorously constrained zones can provide stronger authority containment than a highly fragmented topology joined by broad conduits.
Segmentation should therefore be evaluated by which relationships cease to exist, which authority remains possible across the relationships that remain, and what additional failure dependencies the enforcement mechanisms introduce, rather than by zone count alone.
ZTNA relocates trust; it does not abolish it
Zero Trust Network Access removes the assumption that network membership alone establishes trust, but it does not create a trust-free architecture. A ZTNA implementation can depend on identity providers, certificate authorities, endpoint-posture systems, policy engines, brokers, DNS, endpoint telemetry, cloud services, administrative platforms, and communications infrastructure. These dependencies become part of the same identity, availability, and authority architecture that the ZTNA service is intended to protect.
If remote engineering is operationally important, failure or compromise of that infrastructure can reduce the trusted actions available to operators just as compromise of a traditional VPN can increase attacker reachability. The security improvement lies in making admission resource-specific, identity-specific, endpoint-sensitive, and contextual rather than in eliminating trust.
ZTNA is consequently well suited to human administrative and engineering sessions where identity, device state, session context, target resource, and policy can be evaluated before access is established. Deterministic millisecond-scale controller communication is a different architectural problem and should not be forced through a continuously evaluated cloud authorization path merely to satisfy a literal interpretation of zero trust.
The principle remains the one established earlier: trust should be explicit, narrow, continuously governable where practical, and attached to the specific authority being granted rather than inherited from network location.
Data diodes gain security by removing capability
A unidirectional gateway has an unusually strong security property precisely because one category of communication capability has been removed rather than conditionally permitted. As established earlier, a conduit d can permit OT-to-external transfer while making the reverse flow set empty. Unlike a firewall rule that can potentially be modified, misconfigured, or bypassed through another accepted protocol, a correctly implemented physical unidirectional boundary removes reverse communication through that specific conduit by construction.
The same restriction necessarily removes legitimate functions. Interactive troubleshooting, acknowledgements, request/response APIs, remote control, bidirectional synchronization, and other workflows cannot traverse a strictly one-way path in the prohibited direction. The security property and the functional limitation are the same architectural fact viewed from different objectives. The design question is therefore not whether a diode is generically more secure than a firewall. It is:
Where does the industrial mission permit reverse capability to be removed entirely?
Where the answer is nowhere, a diode is the wrong mechanism; where reverse communication has no mission requirement, retaining it merely for architectural convenience preserves unnecessary capability. This is the least-functionality principle applied at the conduit level.
Cryptography creates a key-management system
Secure industrial protocols materially improve peer authentication, message integrity, and confidentiality, but they also create a second lifecycle system that must remain trustworthy for as long as the industrial function depends on the cryptographic protection. Certificates, trust lists, certificate authorities, device identities, enrollment mechanisms, private keys, renewal, revocation, secure time, algorithm selection, and eventual cryptographic migration all become part of the operating architecture.
A protocol has therefore not been meaningfully modernized merely because TLS was enabled during commissioning. A ten- or twenty-year industrial deployment must answer how device certificates are renewed, what happens when the issuing CA changes, how a credential can be revoked at a disconnected site, whether legacy hardware protects private keys adequately, how trust stores are updated, what happens when cryptographic algorithms become obsolete, and whether the fleet can migrate without simultaneous replacement of every controller.
This is especially important in distributed OT because certificate lifecycle mechanisms can themselves become critical dependencies. A highly secure device identity model that cannot be renewed after the original supplier infrastructure disappears may become operationally weaker than a simpler mechanism with a sustainable lifecycle.
Cryptographic security therefore requires cryptographic lifecycle governability, and crypto agility becomes part of industrial maintainability rather than a purely cryptographic concern.
Brownfield modernization can require replacing a cell, not a device
Security capability in brownfield OT is constrained by compatibility. A secure controller mode may require corresponding HMI software, engineering tools, communications processors, firmware, peer devices, certificates, protocol versions, or gateway functionality. Replacing one controller can therefore break operational relationships that the surrounding cell still requires, even when the replacement device is technically more secure in isolation.
The economically and technically meaningful modernization unit may consequently be the process cell or dependency cluster, not the individual asset. Let \mathfrak{M}_{\mathrm{mod}} denote the set of technically feasible modernization plans. A conceptual planning model is
Here, R_{\mathrm{lifecycle}}(M) is a decision criterion representing the residual lifecycle risk associated with plan M, rather than a universally measurable risk quantity; C(M) and T_{\mathrm{outage}}(M) represent cost and required outage, while the safety and mission predicates require the migration itself to remain compatible with the process.
The model makes explicit why brownfield modernization is a constrained engineering problem rather than a simple product-refresh exercise. A technically superior controller may be infeasible if deploying it requires an unacceptable plant outage, destroys compatibility with safety-certified equipment, or requires simultaneous migration of a surrounding system that the organization cannot yet replace.
Compensating controls are therefore legitimate in brownfield OT when they are explicit, evidence-based, periodically reassessed, tied to known product limitations, and associated with a migration condition under which they can eventually be retired. They become dangerous when an explicitly temporary exception silently becomes permanent architecture.
Deterministic networks are not stationary
The detection architecture developed earlier exploits the relative predictability of industrial communication, but predictability does not imply stationarity. Legitimate behavior changes with startup, shutdown, maintenance, production campaigns, seasonal demand, equipment replacement, failover, degradation, control tuning, and asset aging.
This is why the expected communication model was conditioned on operational state z rather than represented as one universal baseline. Even that is incomplete over long lifecycles: two periods classified as the same nominal production state can legitimately differ after a plant modification, software migration, controller replacement, or operational-policy change.
Statistical behavior can therefore drift from an original distribution \mathcal{D}_0 to a later legitimate distribution \mathcal{D}_t without attack. A detector frozen against the original baseline will accumulate false positives, while a detector that adapts automatically to every persistent change creates the opposite danger: slowly introduced malicious behavior can become part of the learned definition of normality.
The open technical problem is therefore not generic anomaly detection. It is context-aware adaptation whose baseline can evolve with legitimate engineering change without allowing adversarial behavior to normalize itself silently into the model. That requires connection between detection, asset configuration, process state, maintenance records, and change governance rather than purely statistical adaptation.
Statistical normality is not operational legitimacy
Even a perfect behavioral model cannot answer every question of industrial intent. A command can originate from the expected HMI, use the normal protocol, occur at the usual time, contain values frequently observed in the past, and still be unsafe for the physical state that exists now.
Historical frequency therefore provides evidence about expected behavior but not authorization of physical intent. A detector trained only on statistical regularity can miss a dangerous command that is syntactically and behaviorally ordinary while being inconsistent with current process inventory, equipment state, safety conditions, production authorization, or maintenance context.
This is why the semantic-event model developed earlier includes operational context separately from peer, protocol, and operation expectations, and why process-aware detection remains distinct from network anomaly scoring.
Normality describes what has commonly happened; legitimacy asks whether the action should happen under the present physical and organizational conditions.
Process residuals do not identify their own cause
The process-aware detection section used residuals between observed and model-predicted behavior as evidence that the physical process no longer behaves as expected. That residual is valuable precisely because it can expose inconsistency independently of some digital control signals, but it does not identify why the inconsistency exists.
The same residual may be compatible with mechanical failure, sensor drift, calibration error, incorrect model assumptions, an unmodeled disturbance, changed material properties, abnormal but legitimate operating conditions, or cyber manipulation. Treating the residual as an additive mixture of separately observable fault, model, and cyber components would therefore imply identifiability that the detector usually does not possess.
Process residuals provide evidence of inconsistency, not automatic causal attribution. Attribution requires correlation with other evidence layers, identity, engineering activity, network operations, controller state, maintenance context, independent instrumentation, and potentially physical inspection.
This distinction is particularly important where automated response is contemplated. A detector can sometimes justify a bounded protective action because the process is moving outside an acceptable envelope even when the cause remains unknown; it should not claim that a cyberattack has been proven merely because a model residual is large.
Encryption reduces passive semantic visibility
Modern industrial protocol security can reduce what passive monitoring systems can decode. Encryption hides precisely the message semantics that an IDS may previously have inspected, creating a genuine trade-off between communications confidentiality and passive protocol visibility.
The correct response is not to preserve insecure protocol modes merely so that a passive sensor can continue reading traffic. Observability should instead move toward trusted components that can expose the required semantics after decryption or at the point where the operation is generated or executed: controllers, engineering applications, access brokers, protocol gateways, application servers, authenticated event streams, or other trusted endpoints.
This changes the monitoring architecture. Packet inspection becomes one evidence source rather than the universal observation point, while endpoint and application evidence become more important as cryptographic protection increases. The design objective is therefore:
Communication confidentiality and semantic observability must be composed rather than treated as mutually exclusive objectives.
A secure architecture should protect communication on untrusted paths while ensuring that sufficiently trustworthy systems still expose the identity, operation, target, result, and state changes required for detection and investigation.
AI is an analytical accelerator, not an established autonomous safety authority
The August 2026 Siemens advisory provides concrete evidence that attackers are already using AI assistance to generate and iterate industrial exploitation scripts.110 Defenders can likewise use AI for log summarization, alert correlation, protocol-documentation analysis, configuration review, malware triage, forensic reconstruction, and interpretation of unfamiliar engineering artifacts.
These applications can materially reduce analyst workload and accelerate investigation, but they do not establish that present AI systems should independently acquire high-consequence process authority. NIST’s Generative AI Profile identifies risks including confabulation, information-integrity problems, information-security risks, and inappropriate reliance on model output.111 The international Principles for the Secure Integration of Artificial Intelligence in Operational Technology further identifies OT-specific concerns including safety, reliability, data integrity, latency, and new connectivity dependencies, while recommending bounded active control, appropriate oversight, and failsafe paths toward conventional automation or manual operation.112
Figure 13 represents a defensible present authority boundary.
%%{init: {"theme": "neo", "look": "handDrawn", "layout": "elk"}}%%
flowchart TD
A[AI inference or<br/>analytical output]
B[Bounded recommendation<br/>or candidate action]
C{Independent authorization,<br/>operating envelope and<br/>safety constraints}
D[[Authorized industrial<br/>command authority]]
E{{High-consequence controller<br/>or process-control state}}
F(((Physical process<br/>effect)))
A -->|produces| B
B ==>|candidate admitted for execution| D
C -.->|required independent conditions| D
D ==>|authorized command changes state| E
E ==>|control state affects process| F
Figure 13: Bounded AI authority in high-consequence OT. Rectangles represent analytical output and candidate actions; the brace represents independent authorization and safety conditions; double brackets represent privileged industrial command authority; the double-curly node represents high-consequence controller or process-control state; the terminal rounded node represents physical consequence. Dotted edges represent required independent conditions and thick edges represent escalation of authority or consequence; AI output alone does not grant privileged industrial authority.
The independent condition may involve a human operator, deterministic safety logic, conventional control logic, hard operating envelopes, formal interlocks, or a combination of these mechanisms. The architectural point is not that autonomous AI can never become sufficiently assured for higher-consequence use; it is that model output should not silently bypass the authorization and safety mechanisms applied to other industrial decision paths.
AI can therefore be highly useful upstream of authority while still being deliberately bounded at the transition where recommendations become commands.
Recovery assurance remains an open problem
A backup hash proves that an artifact corresponds to a previously recorded digest. It does not prove that the artifact was trustworthy when archived, that the firmware beneath it is trustworthy, that compromised credentials have been revoked, that persistence does not remain elsewhere in the architecture, that safety configuration is correct, or that physical actuator state corresponds to the digital state the recovered controller assumes.
There is therefore no universal cryptographic equivalent of a known-good plant. Cyber-physical recovery assurance must compose several forms of evidence: artifact provenance and integrity, firmware and controller state, identity state, network and security configuration, approved engineering state, safety state, recovery-tool integrity, and physical-process reconciliation.
This is fundamentally harder than ordinary file restoration because the recovered object is not one digital artifact. It is a distributed cyber-physical state whose correctness depends on relationships among software, identities, configuration, controllers, field equipment, human intervention, and the physical process.
The open problem is how to produce strong, repeatable, and sufficiently automated assurance that a reconstituted plant is not merely running again but has returned to a defensible trusted state.
Controller attestation is a promising direction
Hardware-rooted controller attestation could improve one important part of that assurance problem by allowing an independent verifier to obtain fresh cryptographic evidence about security-relevant controller state. Suppose controller c possesses device key k_c protected by an appropriate hardware root of trust. For a fresh challenge nonce n, a conceptual attestation object could be
where F_c represents measured firmware state, C_c the security-relevant controller configuration, P_c the deployed project state, and n provides freshness against replay.
A verifier holding an independently trusted reference and the appropriate device identity could then establish that the controller reports a particular measured digital state. At fleet scale, such a mechanism could materially strengthen integrity assessment, commissioning verification, incident triage, and recovery assurance compared with manual comparison of engineering projects alone.
The engineering difficulties are substantial. A practical design must determine which state is measurable and security-relevant, which runtime changes are legitimate, how the device key is provisioned and protected, how measurement works with redundant controllers and safety-certified platforms, what performance overhead is acceptable, and how the trust architecture survives long industrial lifecycles. It must also support cryptographic migration when the attestation algorithm, CA hierarchy, or hardware root itself becomes obsolete.
Attestation also does not establish physical state. A cryptographically attested controller with verified firmware and project logic can still resume operation against a manually repositioned valve, an unexpected breaker state, or process inventory that changed while automation was unavailable.
Controller attestation would therefore strengthen one layer of the trust argument rather than solve cyber-physical recovery assurance in full.
Industrial semantic authorization remains too coarse
Modern industrial security still exhibits a large gap between authenticated access and authorized physical intent. A sufficiently expressive policy might state that a particular engineer may download independently approved project revision 8.3 to PLC 12 between 09:00 and 11:00 under work order 4711, from one approved engineering workstation, while remaining unable to modify controller protection or safety functions.
Such a policy combines human identity, endpoint identity, target controller, industrial operation, approved artifact, time, change authorization, and safety restrictions. Many brownfield environments cannot enforce that complete relation end to end. Authorization remains fragmented among network ACLs, workstation login, shared or device-local passwords, application roles, change procedures, procedural approval, and operator supervision.
The open problem is therefore not simply stronger authentication. It is how to represent and enforce industrial intent at the semantic level at which physical authority actually exists.
Closing the gap requires finer-grained controller and protocol authorization, stronger linkage between engineering tools and approved artifacts, machine-readable maintenance context, trustworthy endpoint identity, and enforcement mechanisms able to distinguish one legitimate industrial operation from another rather than granting broad authority after successful login.
Cybersecurity and functional safety still use partially separate assurance models
Cybersecurity and functional safety analyze overlapping physical systems from different assurance traditions. Cybersecurity asks which attacker capabilities can be derived, which trust relationships can be abused, and which controls make those derivations infeasible; functional safety asks which hazards exist, which random or systematic failures can produce them, which protective functions reduce the resulting risk, and which common-cause failures invalidate assumed independence. For cyber-physical systems, the stronger combined claim is:
The safety argument remains valid for the defined cyber-compromise assumptions under which the system is intended to operate.
Producing that claim requires tighter composition among threat modeling, capability-inference analysis, IEC 62443 architecture, hazard analysis, safety-integrity analysis, common-cause assessment, degraded-mode analysis, and verification of shared dependencies. It becomes especially important where ordinary control and safety share engineering tools, privileged identities, networks, firmware supply chains, remote-maintenance paths, or management infrastructure.
The unresolved problem is therefore not merely whether the SIS is cybersecure. It is whether the cyber architecture preserves the independence, diagnostic assumptions, response times, and failure properties on which the functional-safety case depends.
A remote PLC or RTU can have modest consequence when considered individually while the fleet has high aggregate consequence if many devices share the same credential, firmware, cloud service, integrator template, telecommunications gateway, software defect, management interface, or administrative control plane.
Let \mathcal{D}_{\mathrm{fleet}} denote the set of security-relevant dependencies shared across the distributed fleet. Compromise of one dependency d\in\mathcal{D}_{\mathrm{fleet}} can then alter the security state of many nominally separate sites simultaneously, even where the sites have no direct lateral connectivity.
Per-device risk assessment can consequently underestimate correlated common-mode exposure. The relevant architectural question becomes:
Which supplier, credential, firmware defect, communications service, cloud platform, integrator configuration, or management plane can create simultaneous authority over a material fraction of the fleet?
This is structurally similar to the shared-dependency problem inside one plant but can operate at much larger scale. Geographic distribution provides little security independence when hundreds or thousands of devices inherit the same digital root of trust.
Table 14 summarizes the principal unresolved engineering problems exposed by the architecture.
Open problem
Why current practice is incomplete
Desired property
Controller attestation
Backups and version checks do not necessarily prove running controller state
Hardware-rooted, fresh, independently verifiable evidence of security-relevant controller state
Semantic authorization
Network and identity authorization remain too coarse for physical intent
Policy over specific industrial operations, targets, artifacts, identities, times, and process contexts
Cyber-safety co-assurance
Safety and cybersecurity cases are often developed partly independently
Safety claims explicitly conditioned on defined cyber-compromise assumptions
Fleet-correlated risk
Device-by-device analysis hides common credentials, firmware, suppliers, services, and control planes
Common-mode dependency analysis across distributed assets
Safe automated containment
Cyber isolation can itself damage availability or safety
Machine-speed containment whose physical consequences are explicitly bounded by design
Encrypted OT observability
Cryptography removes industrial semantics from some passive inspection points
Trusted controller, application, endpoint, broker, or gateway evidence
Adaptive anomaly detection
Static models drift while adaptive models can normalize malicious behavior
Context-sensitive adaptation coupled to trusted change information and resistant to adversarial normalization
Process-state verification
Valid cyber commands can remain physically inappropriate
Diverse evidence linking command, controller state, process state, and physical behavior
Recovery assurance
Known-good digital files do not establish known-good plant state
Evidence-based cyber-physical reconstitution and restart criteria
Cryptographic agility
Industrial assets can outlive algorithms, CAs, protocols, and PKI designs
Long-lived cryptographic migration without wholesale equipment replacement
AI-assisted OT defense
AI can accelerate analysis while producing incorrect or ungrounded conclusions
Bounded, evidence-linked analytical assistance with independent authorization for consequential action
Cross-organizational authority
Operators, carriers, cloud providers, integrators, and vendors can participate in one authority structure
End-to-end assurance across organizational trust boundaries
Table 14: Principal research and engineering gaps exposed by the reference architecture.
Residual risk therefore remains, and the architecture should state explicitly which compromise conditions it is designed to tolerate rather than imply protection against arbitrary attacker authority. A direct compromise of an already high-consequence controller with unrestricted command capability begins beyond several of the defensive gates described in this article; a compromise of an ordinary vendor identity, enterprise endpoint, remote site, or engineering-admission mechanism does not.
For the defined design-basis footholds, the objective is that compromise remains bounded: reachable resources remain constrained to mission-relevant relationships, operational authority remains narrower than the full capability of the compromised system, deviations remain sufficiently observable, physical consequence remains limited by independent control or protection, and trustworthy state remains recoverable after containment.
The resulting engineering principle is therefore:
Compromise should remain bounded rather than becoming automatically composable into uncontrolled physical authority.
That is a more realistic objective than eliminating compromise altogether, and a more demanding one than merely adding security products. It requires an architecture in which reachability, authority, observation, physical consequence, and recovery are deliberately constrained by different and partially independent mechanisms.
From reachable controllers to resilient cyber-physical systems
The official warnings and incidents examined in this article do not reveal an entirely new industrial-security problem; they reveal the maturation, scaling, and operational exploitation of a long-standing architectural one. Industrial systems have always depended on assumptions about who can reach controllers, which engineering systems are trustworthy, which communication relationships are legitimate, which credentials remain controlled, and which people or applications possess authority over the physical process. What has changed is the environment surrounding those assumptions.
Remote maintenance is now routine, geographically distributed assets are connected through carrier and cloud infrastructure, industrial information is consumed by enterprise and external platforms, Internet-wide scanning makes exposed OT discoverable at scale, industrial protocol libraries make legitimate controller functionality easier to automate, and AI-assisted development can further reduce the effort required to convert public technical information into working industrial tooling.113
The decisive question is therefore no longer simply:
Does this PLC contain a vulnerability?
It is:
Can a plausible digital foothold, together with the trust relationships and technical capabilities of the surrounding architecture, derive enough operational authority to create an unacceptable physical consequence?
That question shifts the security object from the individual device to the architecture through which authority is created, constrained, observed, and recovered.
The evidence points to architecture rather than one vendor
The incidents examined in this article span heterogeneous industrial technologies. The 2026 U.S. activity is broader than a single controller family, while the Polish incidents involved combinations of RTUs, protection infrastructure, Windows systems, firewalls, serial devices, cellular gateways, WAGO controllers, and Siemens PLCs. Other cases discussed earlier involved exposed edge infrastructure, water-control environments, remote dam interfaces, and operational dependence on supporting digital services.
The technologies differ, but the recurring structure is the same: an initially limited compromise acquires reachability, reachability exposes additional authority, and that authority becomes security-significant when it can alter controller state or physical behavior. The capability-inference model developed throughout this article makes the important refinement that this progression is not necessarily a simple linear path: high-consequence authority may require several conjunctive prerequisites involving identity, endpoint trust, routing, protocol capability, controller authorization, engineering context, and process state.
The persistent problem is therefore not one manufacturer’s controller design. It is how industrial authority is distributed, inherited, constrained, observed, and recovered across heterogeneous technical and organizational systems.
Public exposure is the easiest dangerous relationship to remove
Direct Internet exposure of PLCs, RTUs, HMIs, engineering services, and industrial edge-management interfaces remains one of the clearest unnecessary sources of reachability. Where no mission requirement exists for an Internet-to-OT conduit, least functionality requires that the corresponding permitted flow set remain empty.
Removing unjustified public exposure is therefore an urgent and comparatively direct risk-reduction measure, but the stronger principle applies to every conduit, including private APNs, MPLS networks, enterprise routes, supplier VPNs, cloud connections, and inter-site links.
For system or principal s, let \mathcal{R}_{\mathrm{allowed}}(s) denote the resources reachable under the deployed architecture and \mathcal{R}_{\mathrm{mission}}(s) those required for the approved mission. The least-reachability objective is
A private carrier network can violate this property just as readily as the public Internet if it creates unnecessary lateral reachability. A vendor VPN can violate it when one authenticated session exposes an entire routed OT network. A compromised management plane can silently redefine it by changing firewall, routing, or identity policy.
Removing public exposure is therefore the first obvious control, not the final architectural objective. The deeper requirement is that every permitted relationship exist because the industrial mission requires it.
The authority surface is the more useful security object
Reachability determines where a principal can communicate; authority determines what that principal can cause after communication succeeds. The distinction is critical because two reachable systems can present radically different consequences: a historian restricted to process-data reads and an engineering workstation with persistent project-download, firmware, and CPU-mode authority are not equivalent security objects.
For the deployed system, let \mathcal{A}_{\mathrm{allowed}} denote the set of operational and administrative actions technically available through legitimate system relationships and \mathcal{A}_{\mathrm{mission}} the actions actually required for approved operation.
Define the excess-authority set as \mathcal{A}_{\mathrm{excess}}=\mathcal{A}_{\mathrm{allowed}}\setminus\mathcal{A}_{\mathrm{mission}}. The least-authority objective is
This authority surface extends beyond controller writes. It includes remote identities, programming tools, firmware functions, CPU operating modes, HMI privileges, network-management infrastructure, supplier services, safety engineering, privileged cloud relationships, private WANs, and recovery administration.
Asset exposure alone is consequently an incomplete security metric. The more useful question is:
What operational authority becomes derivable after reachability succeeds?
That question connects network architecture, identity, use control, protocol semantics, engineering governance, controller integrity, and physical consequence into one security model.
Secure communication is not trusted industrial intent
The protocol-security analysis established that a message can satisfy peer authentication, message protection, and operation-level authorization while still being inappropriate for the physical process. Cryptographic assurance answers whether a peer and message satisfy defined cyber trust conditions; it cannot by itself determine whether executing that operation is correct for the physical state that exists now.
Using the message-acceptance predicate defined earlier, let L_{\mathrm{process}}(m\mid x,z) represent whether operation m is legitimate for physical state x under operational context z. The stronger condition is
A message can therefore be cryptographically authentic, integrity protected, and issued by an authorized principal while remaining physically inappropriate. An operator may possess authority to change a setpoint but issue a value inconsistent with current equipment state; an engineer may possess legitimate programming authority while attempting a project download outside the approved process condition; an authenticated remote command may be technically permitted while violating an operational interlock that exists outside the protocol’s authorization model.
That final process-legitimacy condition is why industrial cybersecurity cannot be separated from process engineering. Cyber authorization determines whether an actor may exercise an operation; process state determines whether exercising it is legitimate now.
Resilience assumes that preventive assumptions will eventually fail
A resilient OT architecture does not assume that every credential remains secret, every supplier remains trustworthy, every firewall remains uncompromised, every engineering endpoint remains clean, or every detector classifies activity correctly. Instead, it asks what happens after one of those assumptions fails and whether the resulting compromise remains bounded.
Can one compromised identity move laterally? Can a remote-access session acquire engineering authority? Can an authenticated principal reach controller-programming functions? Can deployed controller state be independently verified afterward? Does independent protection remain independent? Can the affected path be isolated without creating another process hazard? Can operators retain sufficient local control? Can trustworthy digital state be reconstructed? Can physical and digital state be reconciled before automatic control resumes?
These questions distinguish device hardening from cyber-physical resilience. Prevention seeks to make dangerous derivations infeasible; resilience determines whether the physical mission remains safe, observable, controllable, and recoverable when part of the preventive architecture fails.
A system in which compromise is possible but bounded by independent authority, detection, safety, and recovery mechanisms can be more resilient than one that concentrates all protection on preventing the initial intrusion.
Recovery must restore trust, not merely production
A recovered plant is not merely a plant whose servers, HMIs, and controllers are running again. Recovery has to re-establish trusted identity, privileged administration, security boundaries, engineering systems, controller firmware and projects, supervisory services, safety functions, monitoring, recovery infrastructure, and the physical process state against which restored automation will operate.
The safe-reconstitution model developed earlier captures the two simultaneous requirements: the recovered physical state must again be mission-viable under the trusted capabilities of the recovered cyber-operational state, while restoration of connectivity and authority must not recreate derivations from the defined compromise scenarios to high-consequence capability.
Production can therefore resume technically while trusted mission capability has not yet been restored. Compromised credentials may remain valid, attacker persistence may survive in another administrative domain, a controller project may not correspond to the approved baseline, protection may remain impaired, monitoring may still be absent, or the recovered controller may hold digital assumptions inconsistent with physical reality.
Safe reconstitution consequently requires evidence that compromised trust has been revoked or replaced, the known authority structure has been corrected, controller and engineering state correspond to approved baselines, safety and monitoring functions remain trustworthy, and physical conditions have been reconciled before automatic control resumes. The recovery objective is trustworthy mission capability, not digital uptime.
Many of the most consequential OT security decisions occur before a controller is commissioned. Product capabilities determine whether future architects can implement unique identities, strong authorization, secure communications, meaningful audit evidence, configuration comparison, protected updates, secure remote support, controller-state verification, and trustworthy recovery.
The international Secure by Demand guidance and IEC 62443-4-1/-4-2 move part of this responsibility upstream by treating secure product development and component security capability as lifecycle properties rather than matters that each asset owner should reconstruct independently after procurement.114115116
A product without useful security events constrains future detection; a controller without strong identity or authorization constrains future least-authority design; a supplier without a sustainable vulnerability-management lifecycle constrains future remediation; a mandatory cloud dependency can constrain future isolation; and an undocumented proprietary engineering state can constrain future recovery.
Today’s procurement decision becomes tomorrow’s brownfield constraint. Product selection should therefore be treated as part of security architecture rather than as a preliminary commercial activity that precedes it.
Regulation increasingly converts architectural questions into accountability
NIS2, the Cyber Resilience Act, the EU electricity cybersecurity network code, FERC/NERC CIP, U.S. bulk-power supply-chain measures, Australia’s critical-infrastructure regime, and current UK policy development differ substantially in legal structure and should not be described as one uniform regulatory model. They nevertheless increasingly require organizations, suppliers, and management bodies to demonstrate some combination of risk ownership, lifecycle security, supply-chain governance, incident awareness, continuity, recovery capability, evidence, and accountable decision-making.
The architectural consequence is that previously technical questions increasingly acquire an explicit governance owner. Who accepted this conduit? Who approved persistent supplier authority? Who owns the unsupported controller? Who has verified that recovery works? Who is accountable for the shared dependency between ordinary control and safety? Which evidence demonstrates that a security requirement remains true after years of operational change?
Regulation cannot eliminate residual vulnerability, nor should that be its engineering objective. A more meaningful governance target is to eliminate unowned, unexplained, untested, and indefinitely accepted high-consequence authority.
Governance becomes useful when it preserves traceability from a risk or legal obligation to an engineering property, from that property to an implemented control, and from the control to operational evidence that can be reviewed when the architecture changes.
Main takeaway
The central conclusion of this article is that the security object in modern OT is not the PLC, firewall, VPN, or protocol considered separately; it is the complete cyber-physical authority system through which digital identities and systems acquire the ability to influence physical state.
A PLC can be fully patched and remain dangerously reachable. A firewall can enforce correct rules while being administered through a compromised identity. A private APN can exclude the public Internet while making otherwise unrelated sites mutually reachable. A cryptographically secure protocol can carry an authenticated but physically inappropriate operation. A valid backup can restore a digitally correct controller into an incompatible physical state. An IEC 62443-capable component can operate inside an architecture whose achieved security is weak. A formally compliant organization can retain common-mode dependencies capable of defeating several defensive layers simultaneously.
The durable architectural thesis is therefore expressed by the capability-inference model. For the defined design-basis entry set \mathcal{E}_{\mathrm{DB}}\subseteq V, deployed controls K, and high-consequence set H, the desired containment property is
This is not a claim that arbitrary compromise can never produce physical consequence. It is a design-basis statement: for the ordinary footholds the architecture is explicitly intended to tolerate, compromise should not be sufficient to derive high-consequence process authority. The corresponding engineering principle is simpler:
Reachability must never silently become unrestricted control.
Recommendations
For operators, integrators, asset owners, and suppliers, the preceding analysis leads to a practical priority order:
Remove unjustified reachability first. Eliminate direct Internet exposure of PLCs, RTUs, HMIs, engineering services, and industrial management interfaces wherever no explicit mission requirement exists; apply the same test to private WANs, carrier services, cloud routes, supplier VPNs, and inter-site relationships.
Establish the actual authority architecture, not only the network topology. Inventory identities, privileged endpoints, remote-access paths, engineering tools, controller operations, management planes, supplier services, safety dependencies, monitoring dependencies, and recovery authority so that the organization knows which combinations can derive high-consequence capability.
Replace broad remote network membership with mediated and contracting authority. Require named identities, trusted endpoints, MFA where appropriate, brokered sessions, narrow resource admission, controller-specific authorization, explicit programming windows, and auditable privileged engineering activity.
Separate materially different trust and consequence domains. Treat enterprise systems, the industrial DMZ, privileged engineering, supervisory operations, ordinary process cells, high-consequence control, safety or protection, distributed sites, monitoring, security management, and recovery as distinct authority domains whenever compromise consequences justify the separation.
Constrain controller authority after reachability succeeds. Disable unnecessary services and legacy modes, remove default credentials, use secure protocol capabilities where supported, govern project download, firmware modification and CPU operating-mode changes, protect approved engineering state, and verify deployed controller state against independently trusted baselines.
Detect changes in authority, not merely suspicious packets. Monitor new communication relationships, privileged sessions, engineering actions, controller programming, firmware and protection changes, CPU operating modes, deviations from approved project state, and inconsistencies between digital activity, controller state, and physical process behavior.
Engineer containment and recovery before an incident. Test identity revocation, conduit isolation, loss of external dependencies, manual or local fallback, trusted backup restoration, controller reconstitution, safety validation, digital–physical state reconciliation, controlled restart, and progressive reconnection under realistic degraded conditions.
Move lifecycle requirements upstream. Procure products and services that provide secure defaults, governable identity and authorization, secure-only communication modes where feasible, useful audit evidence, configuration comparison, protected update mechanisms, adequate support periods, controlled remote access, transparent dependencies, and recoverable authoritative state.
These priorities deliberately extend beyond prevention. Removing exposure reduces initial reachability; authority controls prevent permitted reachability from becoming unrestricted industrial power; monitoring provides evidence when assumptions fail; safety and local autonomy constrain physical consequence; recovery prevents destructive compromise from becoming permanent; and procurement and governance determine whether those properties can still be maintained years later.
The final objective is therefore not an unreachable controller, many controllers must remain reachable by legitimate industrial systems. It is a resilient cyber-physical authority architecture in which legitimate connectivity remains possible while high-consequence authority is narrow, contextual, independently constrained, observable, safely containable, and recoverable.
That is the common security property connecting the incidents, standards, technical controls, supply-chain questions, and regulatory obligations examined throughout this article.
National Security Agency, Cybersecurity and Infrastructure Security Agency, Federal Bureau of Investigation, U.S. Department of Energy, & U.S. Environmental Protection Agency. (2026, August 19). Defending against an active threat to Siemens S7 Series PLCs.Cybersecurity Advisory AA26-231A.Official website↩︎
UK National Cyber Security Centre. (2026, August 27). Disruptive cyber activity highlights risk from internet-exposed systems and edge devices.National Cyber Security Centre.Official website↩︎
Lyons, J. (2026, August 24). Iran-linked cyberattack shut down a UK power plant.The Register.Article↩︎
National Security Agency, Cybersecurity and Infrastructure Security Agency, Federal Bureau of Investigation, U.S. Department of Energy, & U.S. Environmental Protection Agency. (2026, August 19). Defending against an active threat to Siemens S7 Series PLCs.Cybersecurity Advisory AA26-231A.Official website↩︎
Stouffer, K., Pease, M., Tang, C., Zimmerman, T., Pillitteri, V., Lightman, S., Hahn, A., Saravia, S., Sherule, A., & Thompson, M. (2023). Guide to Operational Technology (OT) Security.National Institute of Standards and Technology, NIST SP 800-82 Rev. 3.DOI↩︎
Federal Bureau of Investigation & U.S. Environmental Protection Agency. (2026, July 30). Malicious cyber actors targeting water and wastewater sector internet-facing programmable logic controllers, causing operational disruptions.FBI Public Service Announcement.Official website↩︎
Hellenic National Cybersecurity Authority. (2026, August 26). Επείγουσα Ανακοίνωση: Επιθετική δραστηριότητα κατά ελεγκτών (PLC) Siemens S7 [Urgent announcement: Offensive activity targeting Siemens S7 PLCs].Hellenic National Cybersecurity Authority.Official website↩︎
CERT Polska. (2026, August 8). Follow-up report of the December 2025 energy sector incident.CERT Polska.Official website↩︎
Cybersecurity and Infrastructure Security Agency, Federal Bureau of Investigation, National Security Agency, U.S. Environmental Protection Agency, U.S. Department of Energy, & U.S. Cyber Command Cyber National Mission Force. (2026, April 7). Iranian-affiliated cyber actors exploit programmable logic controllers across U.S. critical infrastructure.Joint Cybersecurity Advisory AA26-097A.Official website↩︎
Cybersecurity and Infrastructure Security Agency, Federal Bureau of Investigation, National Security Agency, U.S. Environmental Protection Agency, U.S. Department of Energy, & U.S. Cyber Command Cyber National Mission Force. (2026, April 7). Iranian-affiliated cyber actors exploit programmable logic controllers across U.S. critical infrastructure.Joint Cybersecurity Advisory AA26-097A.Official website↩︎
National Security Agency, Cybersecurity and Infrastructure Security Agency, Federal Bureau of Investigation, U.S. Department of Energy, & U.S. Environmental Protection Agency. (2026, August 19). Defending against an active threat to Siemens S7 Series PLCs.Cybersecurity Advisory AA26-231A.Official website↩︎
Federal Bureau of Investigation & U.S. Environmental Protection Agency. (2026, July 30). Malicious cyber actors targeting water and wastewater sector internet-facing programmable logic controllers, causing operational disruptions.FBI Public Service Announcement.Official website↩︎
Hellenic National Cybersecurity Authority. (2026, August 26). Επείγουσα Ανακοίνωση: Επιθετική δραστηριότητα κατά ελεγκτών (PLC) Siemens S7 [Urgent announcement: Offensive activity targeting Siemens S7 PLCs].Hellenic National Cybersecurity Authority.Official website↩︎
UK National Cyber Security Centre. (2026, August 27). Disruptive cyber activity highlights risk from internet-exposed systems and edge devices.National Cyber Security Centre.Official website↩︎
CERT Polska. (2026, January 30). Energy sector incident report – 29 December 2025.CERT Polska.Official website↩︎
CERT Polska. (2026, January 30). Energy sector incident report – 29 December 2025.CERT Polska.Official website↩︎
CERT Polska. (2026, January 30). Energy sector incident report – 29 December 2025.CERT Polska.Official website↩︎
CERT Polska. (2026, August 8). Follow-up report of the December 2025 energy sector incident.CERT Polska.Official website↩︎
CERT Polska. (2026, August 8). Follow-up report of the December 2025 energy sector incident.CERT Polska.Official website↩︎
CERT Polska. (2026, August 8). Follow-up report of the December 2025 energy sector incident.CERT Polska.Official website↩︎
Cybersecurity and Infrastructure Security Agency, Federal Bureau of Investigation, National Security Agency, U.S. Environmental Protection Agency, U.S. Department of Energy, U.S. Department of Agriculture, U.S. Food and Drug Administration, Multi-State Information Sharing and Analysis Center, Canadian Centre for Cyber Security, & UK National Cyber Security Centre. (2024, May 1). Defending OT operations against ongoing pro-Russia hacktivist activity.Joint fact sheet.PDF↩︎
Cybersecurity and Infrastructure Security Agency, Federal Bureau of Investigation, National Security Agency, & international partners. (2025, December 9). Pro-Russia hacktivists conduct opportunistic attacks against U.S. and global critical infrastructure.Joint Cybersecurity Advisory AA25-343A.Official website↩︎
Centre for Cyber Security. (2025, February). The cyber threat against the Danish water sector.Danish Centre for Cyber Security.PDF↩︎
Danish Defence Intelligence Service. (2025, December 18). Russia is responsible for destructive and disruptive cyber-attacks against Denmark.Danish Defence Intelligence Service.PDF↩︎
Norwegian National Security Authority. (2026, February 6). Risiko 2026.Nasjonal sikkerhetsmyndighet.PDF↩︎
North Carolina State Ports Authority. (2026, August). Cyberattack operational updates affecting Wilmington, Morehead City, and Charlotte terminals.North Carolina Ports.Official website↩︎
CERT Polska. (2026, January 30). Energy sector incident report – 29 December 2025.CERT Polska.Official website↩︎
CERT Polska. (2026, August 8). Follow-up report of the December 2025 energy sector incident.CERT Polska.Official website↩︎
Federal Bureau of Investigation & U.S. Environmental Protection Agency. (2026, July 30). Malicious cyber actors targeting water and wastewater sector internet-facing programmable logic controllers, causing operational disruptions.FBI Public Service Announcement.Official website↩︎
UK National Cyber Security Centre. (2025, September 29). Creating and maintaining a definitive view of your OT architecture.National Cyber Security Centre.Official website↩︎
Stouffer, K., Pease, M., Tang, C., Zimmerman, T., Pillitteri, V., Lightman, S., Hahn, A., Saravia, S., Sherule, A., & Thompson, M. (2023). Guide to Operational Technology (OT) Security.National Institute of Standards and Technology, NIST SP 800-82 Rev. 3.DOI↩︎
Australian Signals Directorate’s Australian Cyber Security Centre, Cybersecurity and Infrastructure Security Agency, National Security Agency, & international partners. (2024, October 2). Principles of operational technology cyber security.Australian Cyber Security Centre.Official website↩︎
Stouffer, K., Pease, M., Tang, C., Zimmerman, T., Pillitteri, V., Lightman, S., Hahn, A., Saravia, S., Sherule, A., & Thompson, M. (2023). Guide to Operational Technology (OT) Security.National Institute of Standards and Technology, NIST SP 800-82 Rev. 3.DOI↩︎
International Society of Automation. (n.d.). ISA/IEC 62443 series of standards.International Society of Automation.Official website↩︎
International Electrotechnical Commission. (2024). IEC 62443-2-1:2024 — Security for industrial automation and control systems — Part 2-1: Security program requirements for IACS asset owners.IEC Webstore.Official website↩︎
International Electrotechnical Commission. (2025). IEC PAS 62443-2-2:2025 — Security for industrial automation and control systems — Part 2-2: IACS security protection scheme.IEC Webstore.Official website↩︎
International Electrotechnical Commission. (2015). IEC TR 62443-2-3:2015 — Security for industrial automation and control systems — Part 2-3: Patch management in the IACS environment.IEC Webstore.Official website↩︎
International Electrotechnical Commission. (2023). IEC 62443-2-4:2023 — Security for industrial automation and control systems — Part 2-4: Security program requirements for IACS service providers.IEC Webstore.Official website↩︎
International Electrotechnical Commission. (2020). IEC 62443-3-2:2020 — Security for industrial automation and control systems — Part 3-2: Security risk assessment for system design.IEC Webstore.Official website↩︎
International Electrotechnical Commission. (2013). IEC 62443-3-3:2013 — Industrial communication networks — Network and system security — Part 3-3: System security requirements and security levels.IEC Webstore.Official website↩︎
International Electrotechnical Commission. (2018). IEC 62443-4-1:2018 — Security for industrial automation and control systems — Part 4-1: Secure product development lifecycle requirements.IEC Webstore.Official website↩︎
International Electrotechnical Commission. (2019). IEC 62443-4-2:2019 — Security for industrial automation and control systems — Part 4-2: Technical security requirements for IACS components.IEC Webstore.Official website↩︎
International Electrotechnical Commission. (2020). IEC 62443-3-2:2020 — Security for industrial automation and control systems — Part 3-2: Security risk assessment for system design.IEC Webstore.Official website↩︎
International Electrotechnical Commission. (2013). IEC 62443-3-3:2013 — Industrial communication networks — Network and system security — Part 3-3: System security requirements and security levels.IEC Webstore.Official website↩︎
International Electrotechnical Commission. (2013). IEC 62443-3-3:2013 — Industrial communication networks — Network and system security — Part 3-3: System security requirements and security levels.IEC Webstore.Official website↩︎
International Electrotechnical Commission. (2013). IEC 62443-3-3:2013 — Industrial communication networks — Network and system security — Part 3-3: System security requirements and security levels.IEC Webstore.Official website↩︎
International Electrotechnical Commission. (2013). IEC 62443-3-3:2013 — Industrial communication networks — Network and system security — Part 3-3: System security requirements and security levels.IEC Webstore.Official website↩︎
UK National Cyber Security Centre. (2026, January 14). Secure connectivity principles for operational technology (OT).National Cyber Security Centre.Official website↩︎
UK National Cyber Security Centre. (2026, January 14). Secure connectivity principles for operational technology (OT).National Cyber Security Centre.Official website↩︎
CERT Polska. (2026, January 30). Energy sector incident report – 29 December 2025.CERT Polska.Official website↩︎
UK National Cyber Security Centre. (2025, October 27). Using PAWs in OT environments.National Cyber Security Centre.Official website↩︎
UK National Cyber Security Centre. (2026, May 27). Zero Trust Network Access (ZTNA).National Cyber Security Centre.Official website↩︎
Tang, C., Marron, J., Saravia, S., Fenimore, P., Stea, B., White, C., & Wiltberger, J. (2026). Cybersecurity for the water and wastewater sector: Build architecture.National Institute of Standards and Technology, NIST SP 1800-45.DOI↩︎
UK National Cyber Security Centre. (2026, January 14). Secure connectivity principles for operational technology (OT).National Cyber Security Centre.Official website↩︎
UK National Cyber Security Centre. (2026, January 14). Secure connectivity principles for operational technology (OT).National Cyber Security Centre.Official website↩︎
UK National Cyber Security Centre. (2026, January 14). Secure connectivity principles for operational technology (OT).National Cyber Security Centre.Official website↩︎
National Security Agency, Cybersecurity and Infrastructure Security Agency, Federal Bureau of Investigation, U.S. Department of Energy, & U.S. Environmental Protection Agency. (2026, August 19). Defending against an active threat to Siemens S7 Series PLCs.Cybersecurity Advisory AA26-231A.Official website↩︎
CERT Polska. (2026, January 30). Energy sector incident report – 29 December 2025.CERT Polska.Official website↩︎
National Security Agency, Cybersecurity and Infrastructure Security Agency, Federal Bureau of Investigation, U.S. Department of Energy, & U.S. Environmental Protection Agency. (2026, August 19). Defending against an active threat to Siemens S7 Series PLCs.Cybersecurity Advisory AA26-231A.Official website↩︎
CERT Polska. (2026, August 8). Follow-up report of the December 2025 energy sector incident.CERT Polska.Official website↩︎
Siemens ProductCERT. (2022, October 11). SSA-568427: Weak key protection vulnerability in SIMATIC S7-1200 and S7-1500 CPU families.Siemens Security Advisory.Official website↩︎
Cybersecurity and Infrastructure Security Agency, Federal Bureau of Investigation, National Security Agency, U.S. Environmental Protection Agency, U.S. Department of Energy, & U.S. Cyber Command Cyber National Mission Force. (2026, April 7). Iranian-affiliated cyber actors exploit programmable logic controllers across U.S. critical infrastructure.Joint Cybersecurity Advisory AA26-097A.Official website↩︎
National Security Agency, Cybersecurity and Infrastructure Security Agency, Federal Bureau of Investigation, U.S. Department of Energy, & U.S. Environmental Protection Agency. (2026, August 19). Defending against an active threat to Siemens S7 Series PLCs.Cybersecurity Advisory AA26-231A.Official website↩︎
ODVA. (n.d.). CIP Security™ for EtherNet/IP™ devices.ODVA Technologies.Official website↩︎
DNP Users Group. (2022). The DNP Secure Session Layer: DNP3 Secure Authentication version 6.DNP Users Group Cybersecurity Task Force.PDF↩︎
OPC Foundation. (2025, October 22). OPC Unified Architecture — Part 2: Security Model (OPC 10000-2, Release 1.05.06).OPC Foundation.Official specification↩︎
Stouffer, K., Pease, M., Tang, C., Zimmerman, T., Pillitteri, V., Lightman, S., Hahn, A., Saravia, S., Sherule, A., & Thompson, M. (2023). Guide to Operational Technology (OT) Security.National Institute of Standards and Technology, NIST SP 800-82 Rev. 3.DOI↩︎
National Security Agency, Cybersecurity and Infrastructure Security Agency, Federal Bureau of Investigation, U.S. Department of Energy, & U.S. Environmental Protection Agency. (2026, August 19). Defending against an active threat to Siemens S7 Series PLCs.Cybersecurity Advisory AA26-231A.Official website↩︎
Stouffer, K., Pease, M., Tang, C., Zimmerman, T., Pillitteri, V., Lightman, S., Hahn, A., Saravia, S., Sherule, A., & Thompson, M. (2023). Guide to Operational Technology (OT) Security.National Institute of Standards and Technology, NIST SP 800-82 Rev. 3.DOI↩︎
UK National Cyber Security Centre. (2026, January 14). Secure connectivity principles for operational technology (OT).National Cyber Security Centre.Official website↩︎
UK National Cyber Security Centre. (2026, January 14). Secure connectivity principles for operational technology (OT).National Cyber Security Centre.Official website↩︎
CERT Polska. (2026, August 8). Follow-up report of the December 2025 energy sector incident.CERT Polska.Official website↩︎
Maysey, T., Powell, M., Steel, J., & Sarvia, S. (2026). OT Backup Quick Start Guide.National Institute of Standards and Technology, NIST SP 1339.DOI↩︎
UK National Cyber Security Centre. (2026, January 14). Secure connectivity principles for operational technology (OT).National Cyber Security Centre.Official website↩︎
Federal Bureau of Investigation & U.S. Environmental Protection Agency. (2026, July 30). Malicious cyber actors targeting water and wastewater sector internet-facing programmable logic controllers, causing operational disruptions.FBI Public Service Announcement.Official website↩︎
CERT Polska. (2026, August 8). Follow-up report of the December 2025 energy sector incident.CERT Polska.Official website↩︎
CERT Polska. (2026, August 8). Follow-up report of the December 2025 energy sector incident.CERT Polska.Official website↩︎
Maysey, T., Powell, M., Steel, J., & Sarvia, S. (2026). OT Backup Quick Start Guide.National Institute of Standards and Technology, NIST SP 1339.DOI↩︎
Australian Signals Directorate’s Australian Cyber Security Centre, Cybersecurity and Infrastructure Security Agency, National Security Agency, & international partners. (2024, October 2). Principles of operational technology cyber security.Australian Cyber Security Centre.Official website↩︎
Cybersecurity and Infrastructure Security Agency, National Security Agency, Federal Bureau of Investigation, U.S. Environmental Protection Agency, Transportation Security Administration, Australian Signals Directorate’s Australian Cyber Security Centre, Canadian Centre for Cyber Security, European Commission Directorate-General for Communications Networks, Content and Technology, German Federal Office for Information Security, Netherlands National Cyber Security Centre, New Zealand National Cyber Security Centre, & UK National Cyber Security Centre. (2025, January 13). Secure by Demand: Priority considerations for operational technology owners and operators when selecting digital products.Joint international guidance.PDF↩︎
Cybersecurity and Infrastructure Security Agency & Federal Bureau of Investigation. (2025, January 17). Product security bad practices.CISA Secure by Design guidance.Official website↩︎
Cybersecurity and Infrastructure Security Agency, National Security Agency, Federal Bureau of Investigation, U.S. Environmental Protection Agency, Transportation Security Administration, Australian Signals Directorate’s Australian Cyber Security Centre, Canadian Centre for Cyber Security, European Commission Directorate-General for Communications Networks, Content and Technology, German Federal Office for Information Security, Netherlands National Cyber Security Centre, New Zealand National Cyber Security Centre, & UK National Cyber Security Centre. (2025, January 13). Secure by Demand: Priority considerations for operational technology owners and operators when selecting digital products.Joint international guidance.PDF↩︎
Cybersecurity and Infrastructure Security Agency, National Security Agency, Federal Bureau of Investigation, U.S. Environmental Protection Agency, Transportation Security Administration, Australian Signals Directorate’s Australian Cyber Security Centre, Canadian Centre for Cyber Security, European Commission Directorate-General for Communications Networks, Content and Technology, German Federal Office for Information Security, Netherlands National Cyber Security Centre, New Zealand National Cyber Security Centre, & UK National Cyber Security Centre. (2025, January 13). Secure by Demand: Priority considerations for operational technology owners and operators when selecting digital products.Joint international guidance.PDF↩︎
Trump, D. J. (2026, August 26). Declaring a national emergency to secure the United States bulk-power system (Executive Order 14420).The White House.Official website↩︎
International Electrotechnical Commission. (2018). IEC 62443-4-1:2018 — Security for industrial automation and control systems — Part 4-1: Secure product development lifecycle requirements.IEC Webstore.Official website↩︎
Souppaya, M., Scarfone, K., & Dodson, D. (2022). Secure Software Development Framework (SSDF) Version 1.1: Recommendations for mitigating the risk of software vulnerabilities.National Institute of Standards and Technology, NIST SP 800-218.DOI↩︎
European Parliament & Council of the European Union. (2024, October 23). Regulation (EU) 2024/2847 on horizontal cybersecurity requirements for products with digital elements (Cyber Resilience Act).Official Journal of the European Union / EUR-Lex.Official text↩︎
National Security Agency, Cybersecurity and Infrastructure Security Agency, & partners. (2026, July 29). 2026 minimum elements for a Software Bill of Materials (SBOM).Cybersecurity Information Sheet.Official website↩︎
Cybersecurity and Infrastructure Security Agency, National Security Agency, Federal Bureau of Investigation, U.S. Environmental Protection Agency, Transportation Security Administration, Australian Signals Directorate’s Australian Cyber Security Centre, Canadian Centre for Cyber Security, European Commission Directorate-General for Communications Networks, Content and Technology, German Federal Office for Information Security, Netherlands National Cyber Security Centre, New Zealand National Cyber Security Centre, & UK National Cyber Security Centre. (2025, January 13). Secure by Demand: Priority considerations for operational technology owners and operators when selecting digital products.Joint international guidance.PDF↩︎
Boyens, J., McWhite, R., & Calloway, L. (2026). NIST Cybersecurity Supply Chain Risk Management: Due Diligence Assessment Quick-Start Guide.National Institute of Standards and Technology, NIST SP 1326.DOI↩︎
European Parliament & Council of the European Union. (2022, December 14). Directive (EU) 2022/2555 on measures for a high common level of cybersecurity across the Union (NIS2 Directive).Official Journal of the European Union / EUR-Lex.Official text↩︎
European Parliament & Council of the European Union. (2022, December 14). Directive (EU) 2022/2555 on measures for a high common level of cybersecurity across the Union (NIS2 Directive).Official Journal of the European Union / EUR-Lex.Official text↩︎
European Parliament & Council of the European Union. (2022, December 14). Directive (EU) 2022/2555 on measures for a high common level of cybersecurity across the Union (NIS2 Directive).Official Journal of the European Union / EUR-Lex.Official text↩︎
European Parliament & Council of the European Union. (2022, December 14). Directive (EU) 2022/2555 on measures for a high common level of cybersecurity across the Union (NIS2 Directive).Official Journal of the European Union / EUR-Lex.Official text↩︎
European Union Agency for Cybersecurity. (2025, June 26). NIS2 technical implementation guidance.ENISA.Official website↩︎
European Commission. (2024, March 11). Commission Delegated Regulation (EU) 2024/1366 establishing a network code on sector-specific rules for cybersecurity aspects of cross-border electricity flows.Official Journal of the European Union / EUR-Lex.Official text↩︎
European Parliament & Council of the European Union. (2024, October 23). Regulation (EU) 2024/2847 on horizontal cybersecurity requirements for products with digital elements (Cyber Resilience Act).Official Journal of the European Union / EUR-Lex.Official text↩︎
European Parliament & Council of the European Union. (2024, October 23). Regulation (EU) 2024/2847 on horizontal cybersecurity requirements for products with digital elements (Cyber Resilience Act).Official Journal of the European Union / EUR-Lex.Official text↩︎
European Parliament & Council of the European Union. (2024, October 23). Regulation (EU) 2024/2847 on horizontal cybersecurity requirements for products with digital elements (Cyber Resilience Act).Official Journal of the European Union / EUR-Lex.Official text↩︎
European Parliament & Council of the European Union. (2024, October 23). Regulation (EU) 2024/2847 on horizontal cybersecurity requirements for products with digital elements (Cyber Resilience Act).Official Journal of the European Union / EUR-Lex.Official text↩︎
Federal Energy Regulatory Commission. (n.d.). Cyber and grid security.Federal Energy Regulatory Commission.Official website↩︎
Federal Energy Regulatory Commission. (2026, March 19). FERC action: New reliability safeguards for American power grid.Federal Energy Regulatory Commission.Official website↩︎
North American Electric Reliability Corporation. (n.d.). CIP-013-3 — Cyber Security — Supply Chain Risk Management.NERC Reliability Standards.Official website↩︎
Trump, D. J. (2026, August 26). Declaring a national emergency to secure the United States bulk-power system (Executive Order 14420).The White House.Official website↩︎
Trump, D. J. (2026, August 26). Declaring a national emergency to secure the United States bulk-power system (Executive Order 14420).The White House.Official website↩︎
Trump, D. J. (2026, August 26). Declaring a national emergency to secure the United States bulk-power system (Executive Order 14420).The White House.Official website↩︎
UK Department for Energy Security and Net Zero & Ofgem. (2026, August 5). Reshaping cyber regulation in downstream gas and electricity: Government response.GOV.UK.Official website↩︎
Parliament of Australia. (2018). Security of Critical Infrastructure Act 2018.Federal Register of Legislation.Official text↩︎
National Security Agency, Cybersecurity and Infrastructure Security Agency, Federal Bureau of Investigation, U.S. Department of Energy, & U.S. Environmental Protection Agency. (2026, August 19). Defending against an active threat to Siemens S7 Series PLCs.Cybersecurity Advisory AA26-231A.Official website↩︎
Autio, C., Schwartz, R., Dunietz, J., Jain, S., Stanley, M., Tabassi, E., Hall, P., & Roberts, K. (2024). Artificial Intelligence Risk Management Framework: Generative Artificial Intelligence Profile.National Institute of Standards and Technology, NIST AI 600-1.DOI↩︎
National Security Agency, Cybersecurity and Infrastructure Security Agency, Australian Signals Directorate’s Australian Cyber Security Centre, & international partners. (2025, December 3). Principles for the secure integration of artificial intelligence in operational technology.Cybersecurity Information Sheet.Official website↩︎
National Security Agency, Cybersecurity and Infrastructure Security Agency, Federal Bureau of Investigation, U.S. Department of Energy, & U.S. Environmental Protection Agency. (2026, August 19). Defending against an active threat to Siemens S7 Series PLCs.Cybersecurity Advisory AA26-231A.Official website↩︎
Cybersecurity and Infrastructure Security Agency, National Security Agency, Federal Bureau of Investigation, U.S. Environmental Protection Agency, Transportation Security Administration, Australian Signals Directorate’s Australian Cyber Security Centre, Canadian Centre for Cyber Security, European Commission Directorate-General for Communications Networks, Content and Technology, German Federal Office for Information Security, Netherlands National Cyber Security Centre, New Zealand National Cyber Security Centre, & UK National Cyber Security Centre. (2025, January 13). Secure by Demand: Priority considerations for operational technology owners and operators when selecting digital products.Joint international guidance.PDF↩︎
International Electrotechnical Commission. (2018). IEC 62443-4-1:2018 — Security for industrial automation and control systems — Part 4-1: Secure product development lifecycle requirements.IEC Webstore.Official website↩︎
International Electrotechnical Commission. (2019). IEC 62443-4-2:2019 — Security for industrial automation and control systems — Part 4-2: Technical security requirements for IACS components.IEC Webstore.Official website↩︎
@online{montano2026,
author = {Montano, Antonio},
title = {Operational {Technology} in the {Crosshairs:} {What} the
2025–2026 {Attacks} {Reveal} {About} {Industrial} {Cyber} {Risk}},
date = {2026-09-05},
url = {https://antomon.github.io/longforms/operational-technology-crosshairs-2025-2026-industrial-cyber-risk/},
langid = {en},
abstract = {The 2025–2026 operational-technology threat landscape
shows a convergence of problems that are often analyzed separately.
Internet-facing PLCs and HMIs are being actively targeted; public
industrial-protocol libraries, scanning services, legitimate
engineering tools, and AI-assisted scripting are reducing the effort
required to interact with controllers; private carrier networks,
supplier access, enterprise dependencies, and shared management
infrastructure create indirect paths into systems that are not
publicly exposed; and recent incidents in the United States, Poland,
Denmark, Norway, the United Kingdom, and other jurisdictions
demonstrate that ordinary digital privileges can become operational
authority or disrupt an industrial mission without requiring novel
ICS malware. The analysis therefore begins by distinguishing
exposure, targeting, confirmed compromise, operational effect, and
attribution, and by separating direct OT manipulation, OT
impairment, and digitally induced operational disruption. The
central argument is that the primary security object in modern OT is
not the individual PLC, firewall, VPN, protocol, or vulnerability,
but the cyber-physical authority architecture through which
identities, endpoints, trust relationships, networks, engineering
systems, protocols, privileges, controller functions, safety
mechanisms, and external dependencies compose. The article develops
a taxonomy of sixteen failure classes across six domains (knowledge
and governance, reachability and trust, identity and authority,
control-path integrity, observability, and safety and resilience),
and formalizes one part of that taxonomy with a typed
capability-inference model. Rather than assuming a simple attack
graph, the model represents conjunctive prerequisites and
architecture-dependent inference rules, derives the capability
closure of plausible initial compromises, distinguishes attack
surface from derived operational authority, identifies
high-consequence derivations and consequence concentration, and
frames defensive-control selection as the problem of making
unacceptable capabilities underivable even under defined
control-loss scenarios. The analysis then extends from cyber
derivability to cyber-physical admissibility. Physical-process
dynamics, mission viability, safe degradation, trusted control
capability, external-dependency loss, safety independence,
containment, and recovery are treated as engineering constraints on
cybersecurity design. IEC 62443 is interpreted as a lifecycle and
system-architecture framework rather than a product checklist: zones
and conduits constrain reachability, Foundational Requirements
constrain different parts of the authority system, and component
security capability must compose with asset-owner governance, secure
product development, system integration, service-provider processes,
patch management, and compensating architecture. From this basis the
article derives practical patterns for industrial DMZs, PAWs,
brokered privileged access, ZTNA, private carrier networks,
microsegmentation, unidirectional gateways, controller least
authority, programming windows, engineering-workstation protection,
project and firmware integrity, and secure industrial protocols
including Modbus Security, CIP Security, DNP3 Secure Authentication,
and OPC UA. Prevention is treated as incomplete. The detection
architecture combines context-sensitive communication baselines,
passive protocol-aware observation, semantic analysis of industrial
operations, controller-state verification, cross-layer identity and
engineering evidence, and process-model residuals while
distinguishing statistical anomaly from operational illegitimacy.
Recovery is correspondingly defined as safe reconstitution rather
than server restoration: containment must preserve physical safety,
backups must include authoritative engineering and security state,
recovery trust must remain sufficiently independent from production,
controller and protection state must be verified, digital and
physical process state must be reconciled, and external connectivity
must be reintroduced only after high-consequence derivations remain
infeasible. The final sections move the analysis upstream and
outward. Secure-by-design procurement, secure defaults,
configuration governability, baseline logging, owner autonomy,
secure development lifecycles, support periods, SBOMs, supplier due
diligence, and cryptographic lifecycle management determine whether
the required security properties remain governable over industrial
lifetimes. NIS2, the EU Cyber Resilience Act, the electricity
cybersecurity network code, NERC/FERC requirements, U.S. bulk-power
supply-chain policy, evolving UK energy regulation, and Australia’s
critical-infrastructure regime are examined as distinct legal
mechanisms that increasingly assign explicit accountability for many
of the same architectural properties. A reference defense-in-depth
architecture is consequently derived around constrained operational
authority rather than arbitrary network levels. Its objective is
deliberately narrower than perfect prevention: plausible ordinary
compromise should remain bounded, high-consequence authority should
require increasingly specific and partially independent conditions,
abnormal authority should be observable, safety and mission
capability should survive defined failures, and trustworthy
operation should be recoverable without reconstructing the original
compromise conditions. The remaining open problems include
common-mode trust, security complexity, brownfield modernization,
cryptographic agility, adaptive detection, encrypted observability,
controller attestation, industrial semantic authorization, AI
authority boundaries, cyber-safety co-assurance, recovery assurance,
and correlated fleet risk.}
}