
The two most capable cyber models available this year both sit behind a gate. Anthropic's Claude Mythos 5.1 comes through the Cyber Verification Program, and OpenAI gates GPT-5.6 Cyber behind Daybreak Red, the top tier of Trusted Access for Cyber. Neither is something a company can simply buy. We run both inside the ELTON pipeline for medical device testing, and this issue covers what that has looked like in practice and what it should change about how a device maker buys testing.
One line up front, since this issue departs from the usual posture here: it is about our own program. I will keep it technical.
Anthropic's model split is worth stating plainly, because the coverage blurs it. Fable 5.1 and Mythos 5.1 are the same model with different safeguards. Fable 5.1 is generally available and refuses penetration testing, exploit generation, and binary vulnerability scanning. Mythos 5.1 allows that work for vetted US organizations under the Cyber Verification Program, which grew out of the invitation-only Project Glasswing preview that seeded Apple, AWS, Microsoft, CrowdStrike, and more than 40 other organizations in April.
OpenAI runs a two-color scheme. Daybreak Blue gives verified defenders GPT-5.6 Sol for triage, code review, and patch validation. Approved organizations can go one tier further to Daybreak Red, which adds GPT-5.6 Cyber for authorized penetration testing, red teaming, and exploit development. GPT-5.6 Cyber shipped on August 10. On OpenAI's internal benchmark of exploit-chain, authentication-bypass, and privilege-escalation tasks it completes 95 percent, against 57.3 percent for GPT-5.5 Cyber and 1.5 percent for the standard model.
Red takes more than a form. OpenAI requires SOC 2 Type II or ISO 27001, single sign-on, MFA, role-based access, usage logging, incident response documentation, company-controlled devices, and since September 1, a hardware security key for each approved person. A company that wants to use these capabilities on behalf of customers goes through a separate review on top of that. OpenAI evaluated the ELTON service and granted Trusted Access for Cyber to protect medical devices, and we run Claude Mythos through Anthropic's verified-defender access. Both companies reserve these models for verified defenders, and a customer uses them through ELTON without obtaining them from Anthropic or OpenAI.
The requirements are the point. Most device makers are not going to stand up a verified cyber program to get a model, and they should not have to, so the results need another route to reach them.
The pipeline was built to be indifferent to which model is plugged into it, because the model's job is bounded. A tool-writing model receives ELTON's engineering inputs, meaning tool specifications, documented protocol behavior, and harness code, and it returns tool selection, discovery code, and exploit code. Customer documentation, code, firmware, findings, and device identifiers are never sent to it, on any lane. Inside the customer's tenant, that data is processed only by open-weight models on ELTON compute, with no vendor account, no telemetry, and no outbound calls. Whether a finding is real is decided by deterministic execution on the code, the firmware, or the physical device. We ask a model to write a tool. We never ask it about you.
So swapping Mythos for GPT-5.6 Cyber, or either one for an open-weight model on our own hardware, is a configuration change. The threat register keeps serving as both the test plan and the coverage record, and verification proceeds as before through L1 on the twin, L2 on code and firmware, and L3 on the running device.
What changes is the ceiling on the reasoning layer, and the two models have different shapes. Mythos is better at reading a large software system, the kind of work where it takes a firmware unpacker's output and reasons about a trust chain that runs from a bootloader into a fleet certificate store. GPT-5.6 Cyber's edge is in writing tools for binary work without source, and in refusing less on the dual-use edges where exploit validation lives. Both produce far more candidate findings than the models we ran a year ago. More candidates means more load on the verification layer, and the better the model gets, the more that layer decides the quality of what a customer receives.
When AI Vulnerability Testing is elected, the customer chooses the lane for the tool-writing model. Frontier through ELTON: Anthropic Claude Mythos and OpenAI Trusted Access for Cyber, run under our access and under the controls both companies require. ELTON open-weight, hosted locally: open-source weights on ELTON compute inside the tenant, so no frontier vendor is in the loop at all. Bring your own: a model the manufacturer already licenses, at its own endpoint, with the harness, the tools, and the tradecraft staying ELTON's.
The boundary is the same on every lane, and so are the coverage record, the verification, and the evidence. The lanes differ in reasoning ceiling, in who is in the loop, and in cost.
If you are evaluating AI pentesting this year, the model name is the least durable part of any offer. Models turn over quarterly, access terms shift with little warning (the hardware key requirement arrived with about three weeks of notice), and an export control can remove a model outright, as on June 13 when Fable 5 and Mythos 5 went dark for every customer and stayed dark for eighteen days. What you are buying is the harness around the model, the coverage record it produces, the verification that turns candidates into evidence, and the operator who holds the credentialed access when the gate moves.
We wrote that up as a buyer's guide, built around the questions procurement and regulatory teams have been asking us this quarter: The 2026 Buyer's Guide to AI Pentesting for Medical Devices.
The gates will tighten. Hardware keys were the first sign, partner-program reviews the second, and I expect the next Anthropic and OpenAI cyber releases to arrive with stricter vetting. The gap that matters is opening between organizations that can run these models under controls and those without that access. For a device maker the practical answer is to get the results through an operator who holds the credentials, and to keep the model as a switch you control. Every output should then be judged by the evidence behind it, since the evidence is the one part of this stack that will look the same a year from now.
---
Sources: Anthropic, "Introducing Claude Fable 5.1 and Claude Mythos 5.1," September 2026. OpenAI Help Center, "OpenAI Daybreak, Trusted Access for Cyber Overview." VentureBeat, "OpenAI launches GPT-5.6-Cyber with reduced refusals, 95% completion on advanced cybersecurity tasks," August 10, 2026. Francis Odum, Lawrence Pingree, and Sean Sosnowski, "The Cybersecurity Implications of Claude Mythos and OpenAI Cyber," Software Analyst Cyber Research, May 11, 2026. Penligent, "GPT-5.4 Cyber vs Claude Mythos, Which Model Fits Cybersecurity Work," April 17, 2026, and "OpenAI Daybreak vs Anthropic Mythos, The Vulnerability Market Splits in Two," May 14, 2026.