In brief: Palo Alto Networks is offering Unit 42 Continuous Frontier AI Defense as an ongoing offensive-security test for web applications, APIs and cloud infrastructure. It combines Claude Mythos 5, GPT-5.6-Cyber and open-weight models with Unit 42 experts. The continuous operating model is new; credible effectiveness data is absent.
Validating attack paths and prioritising remedies
The service repeatedly examines an organisation's attack surface, validates possible paths and ranks them by proven exploitability. Palo Alto lists code-level fixes and virtual patches among its outputs. Reuters says the service will be sold globally as an annual subscription, priced according to the selected mix of OpenAI, Anthropic and open-weight models.
This moves conventional red teaming from time-boxed projects towards continuous testing of changing systems. Palo Alto has not disclosed detection rates, false positives, remediation time or controlled customer comparisons. Its public material also does not specify data location, retention of source code and telemetry, or the exact boundary between agent action and human approval.
Pandorex Analysis: model access is not a security outcome
OpenAI provides the most useful technical context. GPT-5.6-Cyber is trained for authorised vulnerability research and exploit validation with fewer refusals on risky cyber tasks. OpenAI's 95% figure, however, measures whether the model completes such requests, not whether it finds valid flaws. OpenAI also reports that GPT-5.6-Cyber trailed GPT-5.6 Sol on an internal vulnerability-reporting evaluation and was less token-efficient on the standard 300-turn ExploitBench setting. Anthropic describes Mythos 5 as stronger than Opus 5 in offensive cybersecurity, but likewise provides no independent result for this service on the linked product page.
Multiple models may reduce distinct blind spots, while increasing the work needed for permissions, logging, data boundaries and reproducible findings. Buyers should therefore focus on validated findings per test hour, false-positive rate, scope coverage and auditable approvals rather than model count. The announcement establishes a new operational service, not an advantage over human-led red teams or established exposure-management platforms.
