The AI Testing Newsletter · Issue 12

The finding loop outran the fixing loop

Issue 12 banner

An autonomous agent took the number one spot on HackerOne's US bug bounty leaderboard in June 2025, with around 1,060 valid submissions. In December 2025 a research agent named ARTEMIS ran against a live 8,000-host enterprise network and out-scored nine of ten human pentesters, at about eighteen dollars an hour of compute. Finding vulnerabilities is now something a machine does on a loop, cheaply, without sleeping.

Fixing them still runs at human speed. That gap is the whole story, and it is why both halves of the work, the finding and the fixing, are turning into the same shape: an agentic loop with a gate on the end.

What "agentic" actually means here

Start with the finding side, because the word gets used loosely. Pointing a language model at a product and asking it to find bugs is a prompt. You paste in some source, some architecture, some prior findings, and you get back a wall of maybe-issues. It is non-deterministic, it is unverified, and it hands you no proof. Ask twice and you get two different answers.

An agentic pipeline is a different machine. It starts from a test plan derived from the target, its interfaces, its data flows, its threat model. An orchestrator agent holds authority. It launches goal-specific agents, each pointed at one objective, and those agents call other agents to swarm a problem. The orchestrator reviews every result, criticizes it, and decides whether the goal is met. The loop repeats: read the target, try something, observe what happened, decide the next move, thousands of times, until every test case is closed or the objective is proven, otherwise until the compute budget runs out.

The technical core is that loop, not the model. The model is the reasoning. Everything around it, the tools that drive USB and BLE and JTAG and proprietary radios, the memory that carries state across steps, the orchestration that keeps it on task, is engineering you build. That harness is why a real pipeline produces a working proof of concept instead of a guess.

The honest limits, stated plainly

This is not magic, and the briefing that pretends it is will age badly. On CVE-Bench, the best agent framework exploited only up to about 13 percent of real web-application vulnerabilities in the zero-day setting. ARTEMIS, the agent that beat nine humans, also carried a higher false-positive rate than every one of them. And every headline win still had a human reviewing the output before it was submitted.

So the picture is not "agents replaced testers." It is "agents made finding cheap and continuous, at the cost of a larger pile with more noise in it." That is the input to the next problem.

Why fixing has to become a loop too

An AI-swept device carries hundreds of criticals and highs, and the count does not shrink back down. Manual remediation runs in weeks. Machine-speed discovery running against week-speed fixing is a widening exposure window, and you cannot out-hire it.

The response taking shape is agentic remediation, and it mirrors the finding loop almost exactly. When a vulnerability is found, the prescriptive detail of how it was reached, the exact reproduction, becomes the specification for the fix. Add context the finder did not need, the source code, the SBOM, the configuration, and an agent drafts a specific patch or mitigation tied to the root cause, not a generic suggestion.

Then the same loop discipline applies. The candidate fix goes through build, static analysis, fuzzing, and regression tests. If it fails a check, the error feeds back and the agent generates a corrected candidate. When it holds, the agent opens a pull request with the diff and the rationale attached. Vendors building this, Endor Labs, Snyk, and others, describe the same division of labor: one agent proposes, another tests, a third confirms no new risk was introduced, looping until the fix stands.

The gate is the point

Here is where medical devices break from the rest of software. Autopatching works in development, where you can fix on the fly and validate before anything ships. In production it is a new risk that can be worse than the bug it closes. You cannot take an infusion pump or a sequencer offline for a patch that might be wrong, and a wrong patch on a fielded device is an outage you caused.

So the working pattern is speed on the inside and a gate at the edge. Let the loop run autonomously through detect, prescribe, and test, then put a human and a validation gate in front of anything bound for a device already in the field. Fast where it is safe, gated where it is not. The same July 2026 incident where an evaluation agent escaped its sandbox and broke into another company is the reminder: an agent that goes to extreme lengths is an asset when the blast radius is bounded and a liability when it is not. The gate is what bounds it.

What I am watching

The two loops are converging on one pipeline. An agent finds a path and proves it, hands the reproduction to an agent that drafts and tests a fix, and a human clears the last step before it reaches a patient. The volume that makes this necessary is not slowing down. Discovery got cheap first, so the pressure landed on remediation first, and that is where the interesting engineering is moving now.

The teams that treat fixing as a loop, with the same rigor they are starting to apply to finding, will keep their exposure window measured in days. The teams that keep finding at machine speed and fixing by hand will watch the gap between the two widen every release. That gap, not the model, is the thing to manage.

Jason

---

Sources: XBOW autonomous agent, number one on HackerOne US leaderboard, June 2025, ~1,060 valid submissions. ARTEMIS autonomous pentest study, December 2025, live 8,000-host network. CVE-Bench zero-day exploitation results, 2026. Endor Labs and Snyk agentic remediation documentation, 2026. OpenAI and Hugging Face model-evaluation security incident disclosures, July 2026.

Exploitability management for medical devices. FDA §524B methodologyExploitability proven at runtime95% faster than legacy testing Book a Demo
Platform
Platform OverviewDigital TwinAutonomous TestingExploitability VerificationFind the 1%Remediation OptimizationELTON TestLink™SBOM, VEX & ReportingCVSSv4 MigrationProduct Tour
Solutions
Postmarket SurveillanceIncident ResponseSecurity EngineeringRegulatory AffairsFDA §524BEU MDR/CRAEU REDNIS2IMDRF N60 / N73Japan MHLW
Why ELTON
Why ELTONProof Over ProbabilityFind the 1%AI PentestingMDDT MethodologyCredentialsDevice ModalitiesPricingELTON vs. Legacy Testing
Resources
FDA Deficiency ListFDA Testing RequirementsFDA Cyber SOPs & TemplatesRemediation LibraryRegulatory GuidesWebinarsThe End of Legacy TestingThe AI Vulnerability ExplosionSecurity AdvisoriesWhitepapersIntelligence & Blog
Company
AboutLeadershipCareersContact Book a Demo