In brief: US evaluator CAISI rates Z.ai's GLM-5.3 as the strongest publicly released open-weight cyber model tested so far. Substantial gaps remain against leading US models. The cited “four months behind” is a historical index position, not a catch-up forecast.
Four evaluations, not one score
CAISI, housed at NIST, ran models as agents with shell and Python access in the same ReAct harness. Four benchmarks cover finding browser-source vulnerabilities and triggering crashes, turning known V8 bugs into code execution, exploiting real open-source bugs, and discovering undisclosed defects in OSS-Fuzz targets.
GLM-5.3 scored 40.4% on SEC-Bench Pro versus 90.2% for the best tested US model. Results were 61.1% versus 100% on ExploitBench, 9.4% versus 44.4% on ExploitGym, and 7.7% versus 23.2% on CAISI's private OSS-Fuzz set. It substantially improved on the previous Chinese frontier in all four. Public weights also make these capabilities more deployable than an API-only system.
Pandorex Analysis
The NIST results correct two simple narratives. Z.ai used its own CyberGym and ExploitBench results to stress proximity to closed US systems; CAISI confirms a large advance over GLM-5.2, but not broad parity. The gap stays especially wide on longer exploit chains.
CAISI's “about four months” is not a calendar prediction. It comes from an item-response model combining task difficulty with historical release dates. “US frontier best” means the highest CAISI-tested US score per benchmark, not necessarily one model throughout. Unreleased systems are excluded, US safeguards were disabled where possible, and OSS-Fuzz is private.
The defensible conclusion is narrower: open weights reached a new measured high in automated vulnerability discovery and exploit development without matching the tested US frontier. Defenders gain stronger local audit agents. Attackers face a lower barrier without provider filters, but benchmark success does not establish reliable autonomous compromise of untested real systems.
