Get DeafLetter
A weekly selection of signals, vulnerabilities and guides. Critical alerts remain optional.
You can unsubscribe at any time. Privacy policy.
Booz Allen Hamilton tested two frontier AI models in a multi-vendor OT laboratory across eight attack scenarios, demonstrating that these systems can autonomously execute the complete kill chain — from initial access to kinetic effects on physical equipment — without any specialized training on industrial protocols. The tests, conducted on a replicated network with PLCs, HMIs, SCADA, variable frequency drives, sensors, and a collaborative robotic arm, show a compression of attack timelines from hours to minutes: a result that renders defenses based on human detection and security through obscurity of proprietary OT protocols obsolete.
- Two frontier AI models completed every objective in eight attack scenarios in an OT lab, without receiving source code, engineering documents, or OT guidance [SOURCE 1]
- One agent manipulated the SCADA display by taking control of a gateway with pre-authenticated sessions to 14 OT devices, while another identified an ARP misconfiguration on a safety-critical protective relay [SOURCE 1]
- The robotic arm test saw the agent identify protocols, discover APIs, gain administrative access, and move the arm in minutes, also mapping fallback paths [SOURCE 1]
- Booz Allen tested 18 U.S. and Chinese models in a September report: roughly two-thirds achieved initial access without credentials, but Counter AI playbooks reduced autonomous attacker success by over 95% [SOURCE 2]
Eight Scenarios, Zero OT Training: How AI Bypassed Industrial Specialization
Booz Allen Hamilton testers provided the models with no source code, engineering documents, or guidance specific to the operational technology environment. Despite this starting condition, the agents demonstrated general reasoning capabilities on industrial protocols and autonomous adaptation to failures. One episode documented in detail shows the mechanism: the first attack path on an HMI failed due to a version mismatch, so the agent checked active sessions, discovered a different client, located modifiable Jython code, and reconstructed the payload [SOURCE 1].
This sequence — a feedback loop without human intervention — is the core of the autonomy paradox described by Booz Allen. The same capability that makes the agent dangerous, namely the ability to reason like an engineer without being one, is also its structural vulnerability: by manipulating what the agent "sees" and "chooses," Counter AI systems reduced success by 95%.
The lab replicated a stratified industrial network with all components typical of a production environment. Human approval was required before every exploit or physical action, a condition that delimits the experimental context but does not diminish the relevance of the results: the latency of human supervision remains the bottleneck, not the agent's execution.
SCADA Gateway with 14 Pre-Authenticated Sessions: Anatomy of a Compromise
One of the eight scenarios saw the agent compromise a SCADA gateway with live sessions to 14 OT devices across both network zones. The source reports the direct quote generated by the agent: "The gateway holds live sessions to 14 OT devices across BOTH zones — every PLC we've met today, pre-authenticated" [SOURCE 1]. The agent then described the ability to write tags through those sessions, effectively able to "make the operator screens lie" and open a pathway to every controller by impersonating the SCADA server [SOURCE 1].
This scenario is particularly significant because it shows a lateral movement pattern that exploits an existing configuration — pre-authenticated sessions — rather than a zero-day vulnerability. The agent mapped the environment, identified the critical asset, and pivoted through infrastructure that was compromised by design, reducing the kill chain to timescales faster than human response.
Another agent autonomously identified a network misconfiguration on a safety-critical device: a protective relay at address 192.168.25.2 broadcasting ARP for 192.168.25.56, with over 1,000 unanswered queries in 45 minutes. The agent proposed an impersonation attack [SOURCE 1].
Robotic Arm in Minutes: When Kinetic Effects Become Real
The most striking test for physical impact involved a collaborative robotic arm. The sequence documented by the source is: probing robotic protocols, identifying the arm, discovering the API, gaining administrative access, mapping safety zones and movement limits, moving the arm — all in minutes [SOURCE 1]. The agent also mapped fallback paths for the arm [SOURCE 1].
Kyle Miller, vice president of infrastructure cybersecurity at Booz Allen Hamilton, commented that "if an attacker is able to compromise the control of a robotic arm used in a production application, it could cause unexpected arm movement, ignorance of safety limits, or damage to nearby equipment" [SOURCE 1]. Miller added that "some OT organizations are prepared, but many are not" [SOURCE 1].
The speed of execution — "all in minutes" — is the data point that alters the defensive paradigm. Traditional detect-and-response cycles, measured in hours or days, have no margin against agents that perform reconnaissance, compromise, and kinetic effect in a compressed time window.
The September Report: 18 Models and the Boundary of Autonomy
In a separate report dated September 2026, Booz Allen Hamilton tested 18 AI models from U.S. and Chinese providers. Roughly two-thirds of these models achieved initial access to a defended network without credentials [SOURCE 2]. Anthropic's Claude Mythos model autonomously completed the full cyber kill chain in a production enterprise network [SOURCE 2].
However, every frontier API model tested scored zero against real vulnerabilities not intentionally introduced [SOURCE 2]. This data delimits the current perimeter of the threat: models excel on vulnerabilities inserted for testing purposes, but do not yet demonstrate autonomous exploitation capability on zero-day or known real-world defects. Booz Allen estimates that most models could reach full autonomous kill-chain capability within six months [SOURCE 2].
Counter AI playbooks developed by Booz Allen reduced autonomous attacker success by over 95%, a figure that frames the emerging battlefield not as AI versus humans, but as autonomous AI versus defensive AI [SOURCE 2]. Booz Allen launched Vellox Labs Guile as a Counter AI product [SOURCE 2].
"Our testing showed that AI agents can operate with a speed, persistence, and engineering-level precision that may outpace organizations that have not implemented foundational OT cybersecurity practices" — Kyle Miller, VP infrastructure cybersecurity, Booz Allen Hamilton
Why It Matters
The dossier does not specify specific remedial measures for critical infrastructure operators. The source does not detail the actual degree of human supervision during the tests: although "human approval was required" before every exploit or physical action, it is unclear whether this approval was obtained case by case or via blanket authorizations [SOURCE 1].
The source does not identify the two specific models tested in the OT lab: Booz Allen described them only as "latest frontier models from leading providers" [SOURCE 1]. It is unclear whether these two models are included in the 18 from the September report or represent a separate test. The full report "The Offensive Frontier: AI as the Attacker" is not independently verifiable: no direct link to the document was found in the available sources [SOURCE 2].
The brief does not document whether the capabilities demonstrated in the lab have already been reproduced by third parties unaffiliated with Booz Allen Hamilton. No infrastructure overlaps emerge linking the described tests to real incidents or campaigns attributed to known threat actors.
Time as the New Attack Surface
The compression of the kill chain from hours to minutes is not an incremental improvement: it is a paradigm break. Proprietary OT protocols, until now defended by scarce public documentation and the specialization required for their manipulation, lose value as an obstacle when an AI agent can reconstruct their semantics through general reasoning. Security through obscurity becomes useless against an attacker that does not need to know the system it is assaulting in advance.
The field that opens is the one described from Booz Allen's angle: a race between the offensive feedback loop speed and the defensive perceptual manipulation capability. The advantage no longer lies in prior knowledge of the target, but in the speed with which each system — attacker and defender — can adapt to the other's response. This is the crux of industrial cyber warfare: no longer humans versus machines, but machines versus machines, with humans setting the parameters and paying the consequences of errors.
Sources
- https://news.lavx.hu/article/ai-agents-breach-industrial-systems-in-booz-allen-tests-exposing-critical-infrastructure-gaps
- https://industrialcyber.co/reports/booz-allen-ai-models-approaching-autonomous-cyberattack-capability-as-critical-infrastructure-response-windows-narrow/
- https://nvd.nist.gov/vuln/detail/cve-2026-86950
- https://nvd.nist.gov/vuln/detail/cve-2026-88772
- https://support.apple.com/en-us/149226
- https://support.apple.com/en-us/149228
- https://support.apple.com/en-us/149229
Information is based on the cited source and current as of publication.
Sources
Get DeafLetter
A weekly selection of signals, vulnerabilities and guides. Critical alerts remain optional.
You can unsubscribe at any time. Privacy policy.