We have spent more than a decade doing one thing: breaking medical devices. Thousands of assessments. More than 600 penetration tests behind regulatory submissions, for devices that went on to clear. FDA has reviewed our work and our methodology more than once, and we changed both in response to what they asked.
The portfolio runs from proton therapy platforms to implantables you could lose in a pocket. Global manufacturers and first-device startups. Peers have called it the largest volume of medical device pentesting done by a single provider. I cannot verify that, but nobody has shown me a bigger one.
Our reputation was built on old-school pentesting. Human, deep, slow. The kind that finds zero-days no automated tool will. That work still matters, and it is also getting hard to tell apart from cheap offerings that look identical on a proposal and are nothing alike on the bench.
And the ground moved. Pentesting is still required. It is no longer sufficient. What follows is the history of why, and what we watched replace it. We stopped believing this is solvable manually or with spreadsheets years ago. The history explains why we built a machine instead.
A short list of vulnerabilities from a pentest used to be reassuring. Today it reads as incomplete.
Penetration testing is a twenty-plus-year-old practice built for enterprise networks, back when security meant the perimeter. That is why so many of the best testers came up through networking. The objective was narrow and honest: get into the network, demonstrate impact, stop. Success meant finding a way in. It never meant finding every weakness.
Early work lived on exposed services: open ports, weak auth, misconfigured servers, anonymous file shares. Then firewalls got good and testing moved inside. Phishing gave us the internal test: assume a foothold, escalate, move laterally, reach the data.
Tactics evolved for two decades. The output never did. A pentest stayed a time-boxed exercise over a huge scope, built to demonstrate plausible attacker impact. Underneath it sat one quiet assumption: you cannot find everything, so find the worst. For an enterprise, that assumption is fair. For a bounded product you could actually finish testing, it is a design flaw.
Web applications collapsed the enterprise perimeter into a handful of apps and VPN endpoints. For the first time, completeness was achievable: scope was small, deterministic, repeatable. The industry kept defining success as high-impact findings anyway.
Mobile narrowed it further. iOS and Android platform controls pushed testing down to APIs and auth flows. Cloud pushed it into configuration review, where the failure modes are finite and known. Completeness became measurable, in theory. Nobody measured it.
Then enterprises fell in love with red teaming, which narrows the objective to proof of access at any cost. Valuable for resilience. The smallest, least complete output of the entire lineage.
Products, and medical devices especially, are the smallest and most deterministic scope of all. The industry brought them methods built for the largest. Coverage kept shrinking at the exact moment regulators started demanding more of it.
Pentesting is human-driven, and that is both its value and its variance. Five firms on the same target will produce five different finding sets, sometimes with almost no overlap. What gets tested, and how deep, is decided by individual judgment. Pricing sets duration. Duration sets depth. There is a real correlation between cost and quality because senior testers are scarce and expensive, for good reason.
Stretching that labor model to product-grade completeness is economically infeasible, and doubly so postmarket, where testing has to recur. In practice a lot of findings exist because one tester got curious about one corner on one Tuesday.
Buyers and regulators tried to patch the variance with certifications and declared day counts. Weak signals. Some of the best testers I have worked with hold no certifications and produce work no checklist predicts.
None of this is a defect. It is the practice working exactly as designed. It just makes pentesting the wrong primary instrument for proving the posture of a regulated product.
A medical device is a bespoke cyber-physical product with a defined architecture, a short list of interfaces, and a commercial life measured in decades. Unlike an enterprise, the full attack surface can be understood and tested. It is possible. So the question regulators now ask is fair: why are you not doing it?
Traditional pentesting works against that goal. A skilled tester optimizes limited hours toward severe, demonstrable findings. Mediums and lows get deprioritized. Whole areas of scope go unexplored because they probably will not produce impact. That is correct behavior for the original mission and the wrong behavior for a premarket submission. Two equally strong testers will hand you materially different reports, and neither can claim to represent all known vulnerabilities in the product.
This is not a failure of testers. It is a mismatch between what the industry sells and what the regulator needs.
The old interpretation was simple: finish development with few findings, call it success. Today a clean report triggers a different question. Was the testing complete? Reviewers have learned that a low count can mean strong security, shallow testing, or quiet exclusions, and the report cannot tell them which.
A pentest report is a negative-evidence artifact. Only failures get documented. The passes disappear. Many findings could mean thorough testing or a weak product. Few findings could mean a strong product or a weak test. Without test-case-level evidence the outcomes are indistinguishable, and manufacturers are now being asked to justify silence on specific interfaces and components.
The FDA 2025 premarket cybersecurity guidance says it plainly:
Manufacturers should provide details and evidence of testing that demonstrates effective risk control.
Manufacturers should ensure the adequacy of each cybersecurity risk control (e.g., security effectiveness in enforcing the specified security policy, performance for maximum traffic conditions, stability, and reliability, as appropriate).
Cybersecurity controls require testing beyond standard software verification and validation activities to demonstrate the effectiveness of the controls in a proper security context.
The expectation now is a complete inventory of cybersecurity test cases, traced to architecture and controls, executed with pass or fail recorded, plus the full history of what was found and fixed across development. This is not V&V. It is testing the efficacy of your security controls. And a point-in-time PDF cannot carry it: reviewers want a living view of posture tied to a specific release, with the ability to reconstruct how every conclusion was reached.
Human testers did not get less valuable. Their role got sharper. Humans are for the things machines still miss: interaction effects, logic flaws, configuration weirdness, exploit chains that only appear when several conditions stack.
Device testing now draws on SBOM analysis, static analysis, dynamic testing, and configuration review, and those sources produce volume that needs context. Pentesting belongs at the end of that pipeline, applied in small, targeted increments. Micro-penetration tests: verify exploitability for a specific test case, under specific conditions, and record the result. Not the primary deliverable. The last mile of proof.
A vulnerability that was not exploited during a pentest is not invalid. Exploitability is one dimension of prioritization. Completeness and traceability stay mandatory regardless.
Here is the uncomfortable part. Executing, tracking, and reporting that level of evidence by hand is economically infeasible. We know because we tried. Done manually, comprehensive testing costs an order of magnitude more, and the spreadsheet dies the day the second release ships.
Scale requires a different shape: model the product architecture, derive test cases from the model, attach execution evidence to specific components and controls, and track every vulnerability across its lifecycle. Not disposable reports. A system of record that can produce defensible evidence for any release, on any day someone asks.
When we wrote the first version of this article, that system did not exist. So we built it. ELTON runs discovery and test execution as an agentic pipeline through Autonomous Testing, settles exploitability on the running product through Exploitability Verification, and keeps the whole record alive release over release in the lifecycle system of record. The micro-pentest model is how the platform spends human hours now: machines produce the coverage, humans prove the chains. This article turned out to be the requirements document.
Legacy pentesting was never designed to carry a product's security posture, and under current review it visibly cannot. Keep humans for what humans are for: judgment, chains, the last mile of proof. Give completeness to a machine that never gets bored and never bills by the hour. That is the direction FDA is already walking. We stopped waiting for it to arrive.
Start with one device. We build the twin from documentation your quality system already produces, run AI discovery remotely, and show you the graph: the handful to fix, and the evidence for everything else.