MedDevice AI Security Weekly · Issue 18

GLM-5.3 brings end-to-end exploit development to open weights

Issue 18 banner

Anthropic's Frontier Red Team published its assessment of GLM-5.3 on September 29. The report finds that the new open-weight model from Zhipu (Z.ai) builds end-to-end exploits on ExploitBench at close to the rate Claude Mythos Preview did five months ago. Its safeguards failed simple bypasses in Anthropic's testing, and modified copies with the refusals removed were released publicly within days of the weights going public. Before GLM-5.3's weights went public in late August, every model that could do this work kept its safeguards under the lab's control.

Anthropic competes with Zhipu, which is one reason the independent numbers below matter as much as Anthropic's.

What the numbers say

Anthropic ran GLM-5.3 on ExploitBench, which asks a model to turn known bugs in Chrome's V8 engine into working exploits. GLM-5.3 built end-to-end exploits in 50 of 410 attempts, and Mythos Preview in 56. On Anthropic's internal binary exploitation benchmark, built from Google OSS-Fuzz projects, GLM-5.3 achieved a full control-flow hijack in 4 percent of trials and Mythos Preview in 6 percent. Claude Opus 4.6 and GLM-5.2, the previous generation from each lab, scored zero on both.

NIST's Center for AI Standards and Innovation (CAISI) published its own evaluation on September 17. It called GLM-5.3 "the most cyber-capable open-weight model released to date." CAISI puts it about four months behind the US frontier on its aggregate index. Its table puts GLM-5.3 at 40.4 percent on SEC-Bench Pro and 61.1 percent on ExploitBench, against 90.2 and 100 percent for the best US model. CAISI's ExploitBench figure looks higher than Anthropic's because CAISI scores each task on a 16-point scale with credit for partial progress, while Anthropic counts only finished exploits. Both reports rank the models in the same order. GLM-5.3 improves sharply on GLM-5.2 and comes close to Mythos Preview on end-to-end exploitation. It still trails the current US frontier, which includes models released only to vetted users.

CAISI scored the US frontier models with cyber safeguards disabled, and several of those models are available only to vetted users. Anyone can download GLM-5.3, and the report shows it can already build some N-day exploits.

The hands-on sessions

The report includes two hands-on sessions, each a day or less with under an hour of human attention. In one, GLM-5.3 found several previously unknown vulnerabilities in the JavaScript engine of a popular browser and produced a working exploit, which Anthropic disclosed to the maintainer. In the same session, the researcher used GLM-5.3 to identify what Anthropic calls exploitable vulnerabilities in wireless and graphics drivers and in network-facing device software. Anthropic is still reviewing those for disclosure. In the other session, the smaller GLM-5.3-Flash took a recently patched Chrome bug from public fix to working exploit in about eight hours of model time. Anthropic priced the compute at $20.40 at Zhipu's API rates.

For comparison, Anthropic's June report on N-day exploitation put the cost of a Windows kernel exploit built with Mythos Preview, then limited to Project Glasswing partners, at roughly $2,000. The bugs differ, so treat the two figures as a rough guide to the cost curve. The weights behind the cheaper one are a public download.

Anthropic's exploit development tasks never triggered a refusal from the released model. The refusals it does have, for overtly harmful requests, fell to simple bypasses in the simulated tests. Anthropic reports that a small team can remove them from the weights at modest cost. Anthropic says the same bypasses failed against its own Claude models, and outsiders cannot edit a closed model's weights.

Access programs and public weights

Zhipu released GLM-5.3 on August 14 and held the weights for two weeks for what it described as safety evaluation with vetted security partners. Reuters reported at launch that the company planned to limit sensitive cyber functions to verified users through a trusted access program of its own. Then the weights went public, and anyone with a copy can use the capabilities that program was meant to gate.

Open weights also cut off one source of visibility into AI-enabled attacks. The reports Anthropic and other US labs have published this year on attackers misusing their models came from activity on the labs' own platforms. Every session there runs under an account. A copy of GLM-5.3 on private hardware leaves no such record, so attackers who switch to it will drop out of those reports.

For device makers

A published fix for a component in your device can now serve as an exploit specification for a model anyone can download. Medical device patches often take months to reach the field, because each fix needs validation and regulatory coordination before hospitals can deploy it. Anthropic's June report already showed N-day exploits built in hours, and GLM-5.3 brings that speed to attackers who could never pass a vetting review.

Infusion pumps and bedside monitors depend on the same classes of component as the wireless drivers and network-facing software found in the browser session. Some device front ends also embed Chromium, which carries the V8 engine used in ExploitBench.

GLM-5.3 gives anyone testing devices with AI a public reference point for what a downloaded model can do.

What I am watching

Anthropic's post asks governments to test whatever comes next, and GLM-5.3 arrived only two months after GLM-5.2. CAISI turned its assessment around in about a month, which is a good sign for that request.

The other open-weight labs are the next question. Anthropic's charts include Kimi K3 and DeepSeek V4.1-Flash, both well below GLM-5.3 today, and the jump from GLM-5.2 shows how fast that can change. I will read the next Kimi and DeepSeek releases closely.

Last, I am watching whether any regulator ties open-weight exploit capability to device patch timelines. Section 524B requires a plan to address postmarket vulnerabilities in a reasonable time, and out-of-cycle patches as soon as possible for critical vulnerabilities that could cause uncontrolled risks. I would like FDA to say whether a reasonable time changes when a public fix can become a working exploit within a day.

---

Sources: Anthropic Frontier Red Team, "GLM-5.3 and the spread of advanced cyber capabilities," September 29, 2026. NIST Center for AI Standards and Innovation, "CAISI's Assessment of Z.ai's GLM-5.3 Cyber Capabilities," September 17, 2026. Anthropic, "Measuring LLMs' Impact on N-day Exploits," June 8, 2026. Reuters, "China's Z.ai says new model nears Anthropic's Mythos 5 in cyber defence tests," August 14, 2026. The Batch, "GLM-5.3 Makes Cybersecurity Gains," August 28, 2026.

Jason Sinchak
CEO, ELTON
Exploitability management for medical devices. FDA §524B methodologyExploitability proven at runtime95% faster than legacy testing Book a Demo →
Platform
OverviewAvoid FDA DeficienciesAvoid Consulting FeesDigital Twin TraceabilityAI MedDevice PentestingExploitability VerificationVulnerability ChainingRemediation OptimizationRemote TestLink™Incident ResponseAutomated VEX & MetricsCVSSv4 Migration
Solutions
EnterpriseStartups / SMBs Postmarket SurveillanceIncident ResponseSecurity EngineeringRegulatory AffairsFDA §524BEU MDR/CRAEU REDNIS2IMDRF N60 / N73Japan MHLW
Why ELTON
Proof over ProbabilityExploitability VerificationFDA MDDTCVSSv4 MigrationQMSR Audit Compliance One SolutionSubscription TestingAI-NativeELTON vs. Legacy TestingThreat-Led AI PentestingCredentialsDevice ModalitiesPricing
Resources
FDA Deficiency ListFDA Testing RequirementsFDA Cyber SOPs & TemplatesRemediation LibraryRegulatory GuidesWebinarsAI NewsletterThe End of Legacy TestingThe AI Vulnerability ExplosionAI Inside the ProductCybersecurity TestingSecurity AdvisoriesWhitepapersIntelligence & Blog
Company
AboutLeadershipCareersPartnershipsData SecurityContact Meet ELTON →