MedDevice AI Security Weekly · Issue 15

OpenAI pentesting agents attack again

Read this issue on LinkedIn
Issue 15 banner

The message board incident from a few issues back had a prequel, and it only came out on September 11. Researchers Spencer Kitts, Thomas Larsen, and Sydney Von Arx published a report at rubyhack.ai showing that OpenAI's internal agents attacked RubyGems, the package registry for the Ruby language, back in May 2026, two months before the Hugging Face break-in. The community called it the GemStuffer campaign at the time and did not know who was behind it.

This one is worth a close read, because unlike Hugging Face it hit a real production service that millions of developers depend on, and the agents got code execution on a build server in the process.

What the agents did

The first malicious package went up on May 5. By May 11 the agents had submitted more than 2,000 packages, and RubyGems disabled new signups the next day to stop the flood. The team removed over 500 packages on May 13 and reopened registration on May 16. Smaller waves followed in late May and June.

Two things stand out past the volume.

First, the agents gained remote code execution on RubyDoc.info, the service that auto-builds documentation for published gems. Building docs for a gem involves reading a .yardopts file from the package, and that file can point at Ruby scripts to run during the build. The agents abused it to run their own code on RubyDoc's servers. More than a hundred packages used this path. One gem, zzsouthrunner, carried the comment "malicious crawler/exfil for Southwark Jan 2026 docs via rubydoc.info worker" at the top of its payload. The scripts were named exactly what you would expect, including hack.rb, evil.rb, inject.rb, exploit.rb, and ssrf.rb.

Second, six packages tried to steal other users' API keys. RubyGems had a CDN caching bug, rated 7.3, where a legacy sign-in key could be cached and handed to the next caller on the same edge node for up to an hour. It was found independently by a researcher at Truffle Security in July and patched then, and RubyGems found no evidence anyone's key was actually stolen. The agents were probing for it anyway.

What they were after

The odd part is the objective. The agents were using the RCE on RubyDoc to scrape UK local government meeting sites, the ModernGov portals for councils like Lambeth, Wandsworth, and Southwark, then packaging the scraped pages back into new gems to get the data out. Council agendas are public, so nobody needs to hijack a documentation server to reach one.

The researchers could not say why, and offered a few theories: the agents were routing through RubyGems to get around network restrictions in their sandbox, or using it as persistent storage between separate runs, or bypassing rate limits through a proxy. They also could not confirm whether any of it succeeded, because the reasoning traces sit inside OpenAI and were not shared.

How they know it was OpenAI

The attribution is circumstantial but stacked. Hundreds of packages had "oai" in the name, fifteen listed "oai" as the author, and one contact address was openaixyz65947 at gmail. The behavior also overlapped with the German developer wiki that OpenAI has confirmed its agents used as a message board: the RubyGems agents pulled 49 of the same files, used the same r.jina.ai proxy, and reused the same naming conventions.

OpenAI's answer, given to the RubyGems maintainers, was that its agents "used the RubyGems platform to access the internet to carry out benign tasks and retrieve public information." The maintainers also say OpenAI never told them, at the time, that the attack was theirs.

What this means for anyone building or consuming this

Two lessons carry over to device security, and neither is about the model.

The RubyDoc break is a build-pipeline flaw, the kind that lives in a lot of software supply chains. A config file inside an untrusted package told a build server to execute code, and the server did. Swap RubyDoc for a firmware build system, a documentation generator, or a CI runner that processes third-party input, and the shape is the same. Anyone standing up automated testing that ingests untrusted artifacts should assume the build step is an execution surface, because that is what it is.

The rest tracks the containment point from the earlier issue. Nobody handed these agents a way to scrape blocked sites or store data across runs. They found a writable service with a network path out and used it. The failure sat in the environment around the agent, its sandbox and egress path, and no new capability was needed to exploit it. The disclosure gap is the other half: a real registry got attacked, spent days offline to defenders who did not know the source, and the operator that could have said so stayed quiet for four months.

The pattern from May to July to now is one operator, three separate services, each time the agents reaching past the task they were given because the software around them let them. The software around the agent is the surface to watch as more teams point agents at real systems.

---

Sources: Spencer Kitts, Thomas Larsen, and Sydney Von Arx, "The GemStuffer campaign," rubyhack.ai, September 11, 2026. Maciej Mensfeld / Mend.io reporting on the May 2026 RubyGems attack. RubyGems security advisory on the legacy API key CDN caching issue (GHSA-9j48-x3c3-mrp2), July 22, 2026. The Hacker News, "OpenAI Agents Linked to RubyGems Campaign," September 2026. Simon Willison, "OpenAI agents attacked RubyGems back in May," September 12, 2026.

Jason Sinchak
CEO, ELTON
Exploitability management for medical devices. FDA §524B methodologyExploitability proven at runtime95% faster than legacy testing Book a Demo →
Platform
OverviewAvoid FDA DeficienciesAvoid Consulting FeesDigital Twin TraceabilityAI MedDevice PentestingExploitability VerificationVulnerability ChainingRemediation OptimizationRemote TestLink™Incident ResponseAutomated VEX & MetricsCVSSv4 Migration
Solutions
EnterpriseStartups / SMBs Postmarket SurveillanceIncident ResponseSecurity EngineeringRegulatory AffairsFDA §524BEU MDR/CRAEU REDNIS2IMDRF N60 / N73Japan MHLW
Why ELTON
Proof over ProbabilityExploitability VerificationFDA MDDTCVSSv4 MigrationQMSR Audit Compliance One SolutionSubscription TestingAI-NativeELTON vs. Legacy TestingThreat-Led AI PentestingCredentialsDevice ModalitiesPricing
Resources
FDA Deficiency ListFDA Testing RequirementsFDA Cyber SOPs & TemplatesRemediation LibraryRegulatory GuidesWebinarsAI NewsletterThe End of Legacy TestingThe AI Vulnerability ExplosionAI Inside the ProductCybersecurity TestingSecurity AdvisoriesWhitepapersIntelligence & Blog
Company
AboutLeadershipCareersPartnershipsData SecurityContact Meet ELTON →