The Protected Machine: AI for Cyber-Physical Security

Z

ZharfAI Team

May 3, 2026Updated July 30, 202610 min read
The Protected Machine: AI for Cyber-Physical Security

In operational technology, a cyber event can stop production, damage equipment, pollute a process, interrupt an essential service, or endanger people. AI can help defenders organize inventories, correlate weak signals, and prioritize investigation. It also introduces software, data dependencies, remote connections, and failure modes into an environment where availability and deterministic control may matter more than analytical novelty.

The design boundary must be unambiguous: AI monitoring is not autonomous safety control. A model may recommend containment or identify an abnormal sequence. It must not bypass a safety instrumented function, rewrite a controller, trip a process, restore service, or suppress an alarm without the specifically engineered and authorized mechanism for that action. Safety remains independent, deterministic where required, tested, and owned by qualified operations and safety personnel.

1. Govern cyber risk inside the physical mission

Define the service and physical consequence before selecting a detector. What must remain safe, available, within environmental limits, and recoverable? Which loss scenarios can injure people, damage assets, reduce product quality, or interrupt customers? Join cybersecurity, engineering, operations, safety, maintenance, quality, incident response, and leadership in the threat and hazard review.

NIST SP 800-82 Rev. 3 is final guidance for operational-technology security and emphasizes OT’s distinctive performance, reliability, and safety requirements. Its page also notes subsequent pre-draft revision activity; that does not turn an early call for input into final requirements. Organizations must map the guidance to their sector, regulation, architecture, and risk.

2. Establish an authoritative asset and dependency model

Inventory controllers, safety systems, human-machine interfaces, engineering workstations, historians, sensors, actuators, remote terminal units, network equipment, time sources, wireless links, vendor appliances, virtualization, and cloud or enterprise dependencies. Assign stable identifiers and record site, zone, owner, function, criticality, firmware, configuration baseline, support status, and last verification.

Passive discovery can reveal communications without injecting traffic into fragile devices, but it is incomplete when equipment is offline or silent. Active scanning may be unsafe unless tested and approved for the device and process state. Reconcile passive observation, configuration databases, drawings, backup repositories, maintenance records, procurement, and physical walkdowns. Keep “observed,” “declared,” and “verified” as separate states.

3. Model zones, conduits, and trust boundaries

Represent not only devices but permitted flows: source, destination, protocol, port or service, direction, frequency, operational purpose, owner, and approved time window. Connect enterprise, demilitarized, control, safety, remote-access, vendor, and field zones. Record unidirectional gateways, serial links, temporary modems, and maintenance laptops that ordinary IT diagrams miss.

The ISA/IEC 62443 series describes a lifecycle and shared responsibilities for industrial automation and control-system security across asset owners, product suppliers, integrators, and service providers. It is a series of consensus standards with parts and editions, not one universal product label. Teams should identify the exact parts and versions required by contracts or programs.

4. Give telemetry provenance and time semantics

Useful evidence may include network flow, protocol commands, controller diagnostics, configuration changes, Windows and Linux events, remote-access logs, authentication, historian values, alarm state, work orders, and physical process variables. For every feed, record collector, clock source, units, sample interval, loss, buffering, transformation, retention, and security boundary.

Clock drift can reverse an incident timeline. A historian may compress values; a gateway may translate addresses; a controller may report local time; a packet sensor may miss encrypted payloads. Preserve raw evidence where proportionate, keep transformation lineage, and expose gaps. A model should say “telemetry unavailable” rather than infer normality from silence.

5. Detect process-aware anomalies

Signature and allow-list controls remain valuable. AI can add sequence, rate-of-change, multivariate, and peer-group detection: a write command at an unusual phase, a valve state inconsistent with flow, or a workstation contacting an unfamiliar destination. Build baselines by operating mode—startup, shutdown, cleaning, batch change, maintenance, emergency—not from one average “normal.”

Use engineering constraints and first-principles relationships where possible. A purely statistical detector may treat a rare but approved safety test as malicious or learn a slowly degraded process as normal. Show which variables and events contributed, the relevant mode, nearest known precedent, confidence limits, and raw evidence. AI for cybersecurity defense covers complementary enterprise detection patterns.

6. Treat model uncertainty as an operating condition

Evaluate missed attacks and nuisance alerts, but also coverage: which assets, protocols, modes, and sites have enough data for the claimed detection? Test maintenance, failover, sensor replacement, calibration, seasonal load, low-volume facilities, and communications loss. Monitor drift in firmware, topology, production recipes, and tag definitions.

If the feature pipeline is stale or an asset moves outside the validated envelope, downgrade the output and alert the operator. Do not allow a confidence score to hide unknown coverage. Maintain deterministic alerts for critical conditions and a manual analysis route when the AI stack is unavailable.

7. Keep response human-led and runbook-bound

An alert should open a case with affected asset, physical function, evidence timeline, likely consequence, uncertainty, and approved response options. The shift operator validates process state; cybersecurity investigates indicators; engineering assesses equipment behavior; safety personnel judge hazard controls; incident command authorizes coordinated action.

Actions such as isolating a workstation, disabling remote access, blocking an enterprise conduit, or moving a process to a safe state have different physical risks. Preauthorize only narrow, tested actions with clear prerequisites and reversal. Require dual control for consequential steps. Log who saw the recommendation, who decided, what changed, and how the plant state was verified.

8. Never let AI bypass engineered safety

Safety instrumented systems, hardwired interlocks, relief devices, protective relays, emergency shutdown, and operator procedures are designed against hazards through specific engineering and assurance practices. A probabilistic anomaly score does not replace them. Keep safety logic and cybersecurity analytics appropriately segregated; prohibit model write access to safety controllers.

Any proposal to connect monitoring with control requires a formal hazard analysis, security risk assessment, management of change, independent verification, testing in representative conditions, fail-safe behavior, and approval by competent authorities. The safe response may be to observe and call a trained operator. “Autonomous” is not a benefit when the action path has not been engineered for the consequence.

9. Reduce attack paths before adding analytics

Remove unused services and default accounts, enforce unique identities where devices support them, segment zones, restrict administrative paths, control removable media, and broker vendor access through approved, monitored sessions. Use multifactor authentication at remote-access boundaries and time-bound authorization. Protect engineering project files, keys, backups, and configuration tools.

The CISA Cross-Sector Cybersecurity Performance Goals provide a voluntary, prioritized baseline of high-impact practices. CISA describes them as a starting point, not an exhaustive program, compliance scheme, or certification. Sector requirements and site hazards can demand more. Use the goals to find basic control gaps before expecting a model to compensate for them.

10. Secure the AI and data pipeline

Place collectors so they cannot become unintended bridges between security zones. Prefer one-way or brokered movement where architecture permits. Authenticate and sign sensor-to-platform communications, validate schemas, restrict model egress, and isolate training from production. Treat protocol descriptions, vendor advisories, tickets, and retrieved documents as untrusted content that can contain prompt injection or malicious instructions.

Pin model, feature, rule, prompt, and retrieval-corpus versions. Approve every tool the model can call and deny writes by default. Keep secrets and raw credentials out of prompts and traces. Test whether a compromised low-criticality sensor can poison a baseline or flood the queue. Recovery must include clean models and feature state, not only clean servers.

11. Prepare for degraded and manual operation

Document minimum safe staffing, local control, communications alternatives, trusted drawings, spare parts, configuration backups, and procedures for loss of enterprise services, identity systems, cloud analytics, GPS time, or remote vendors. Test restoration order because starting an application before authoritative tags, time, or identity can create misleading data.

Backups need offline or otherwise protected copies, integrity checks, known device and firmware compatibility, and practiced restore. Do not assume a controller backup captures field state, calibration, recipes, safety proof tests, or vendor license material. AI and critical-infrastructure risk management places these technical controls in a wider service-continuity context.

12. Exercise attacks without endangering operations

Use tabletop exercises, offline replicas, cyber ranges, digital twins with explicit validity limits, and scheduled site tests. Keep test traffic and credentials isolated from production unless an approved procedure says otherwise. Scenarios should include stolen vendor access, engineering-workstation compromise, false sensor data, loss of view, ransomware in supporting IT, time manipulation, insider action, and simultaneous physical disruption.

The joint CISA and international-partner guidance on principles of operational-technology cybersecurity presents high-level best practices such as making safety paramount and understanding the OT environment. It is guidance, not legislation or a substitute for sector engineering standards. Use it to challenge whether an exercise protects the physical mission.

13. Manage suppliers and change as part of defense

Require asset and software inventories, vulnerability disclosure, secure update processes, support lifetimes, remote-access controls, backup and restore documentation, incident notification, and evidence needed for site assurance. Verify how a vendor model is trained, updated, isolated, logged, and retired. Contractual claims such as “air-gapped” or “no data retained” need architectural and operational verification.

Run management of change for firmware, controller logic, network routes, sensor mappings, model versions, feature definitions, and thresholds. Capture the reason, risk review, test, approvers, implementation window, and rollback. An unrecorded maintenance change can look exactly like an intrusion and can also blind a detector.

14. Measure security, safety, and operational burden together

Relevant indicators include:

  • verified asset and permitted-flow coverage by criticality;
  • telemetry freshness, clock offset, packet loss, and unmonitored intervals;
  • alert precision, detection recall in exercises, and time to qualified triage;
  • safety- or production-sensitive false positives and missed events;
  • remote-access sessions with owner, purpose, recording, and timely closure;
  • unauthorized change attempts and baseline deviations resolved;
  • restore success, recovery time, and manual-operation exercise results;
  • model, rule, and feature changes with completed validation;
  • backlog age by consequence, not only alert count;
  • incidents in which AI advice was rejected, why, and the observed outcome.

Avoid optimizing mean time to containment without measuring physical consequence. A fast network block that removes operator visibility can be worse than a slower coordinated response.

15. Roll out in reversible layers

First verify assets, flows, time, backups, remote access, and existing deterministic detection. Next deploy passive collection and run AI in shadow mode, comparing results with operator and incident records. Then provide explainable alerts to a jointly staffed review queue. Pilot a small set of preplanned response recommendations in exercises before allowing any workflow automation.

Do not connect analytics to control because a dashboard phase succeeded. Any action capability needs its own engineering safety case, authority, validation, failure testing, and rollback. Operational-readiness checks for AI can structure each gate. Expansion should depend on coverage and control evidence, not a promised percentage reduction in analyst work.

Source notes

Sources were reviewed on July 30, 2026. NIST SP 800-82 Rev. 3 was the cited final OT-security publication; pre-draft activity for a later revision is not final guidance. The CISA performance goals and joint OT principles are voluntary guidance, not exhaustive sector rules or certification. ISA/IEC 62443 is a multi-part, evolving standards series; applicability depends on the exact part, edition, role, contract, and jurisdiction. Operators must also follow current safety, environmental, sector, labor, incident-reporting, and national-security requirements. Qualified control, safety, cybersecurity, and operations professionals should approve any change that could affect a physical process.

#Cyber-Physical Security#Critical Infrastructure#Industrial Control#Anomaly Detection#AI

Related Posts

Keep reading

See the daily briefing and the operational guides. This page is an archive note, not an invitation to start a project.