FDA names seven cybersecurity threats for AI devices. ELTON tests all seven, from the training data to the deployed device.
Section XII lists seven cybersecurity threats that are specific to AI devices.
Cyber devices under section 524B carry a cybersecurity requirement already: device software, able to connect, with characteristics that could be attacked. For those devices, the draft guidance adds the AI-specific risks their submissions should address.
Data poisoning, model inversion or stealing, model evasion, data leakage, overfitting, model bias, and performance drift. Six of the seven reach the model through the data and infrastructure around it, not the deployed model alone.
Cybersecurity risk management that covers these AI risks, and testing matched to them, on top of the 2023 premarket cybersecurity guidance. The evidence is what an FDA reviewer reads.
Most AI security testing checks the deployed model. That is one threat of seven.
Hands-on penetration testing across the pipeline: the training data stores, the model registry, the inference API, the runtime inputs, the pipeline traffic, and the postmarket monitoring. Every finding is reproducible and comes with evidence, and it needs no model access from you.
A bounded evaluation of an advanced adversarial technique against the model and its pipeline. One target, one batch, one feed, scored and rated, with a recommended fix. Where full proof is a research project, ELTON bounds it rather than skipping it.
Each activity is tied to one of the seven FDA threats. The result is an AI-Enabled Device Cybersecurity Assessment Report that an FDA reviewer can read against the named risks, with traceable evidence for each.
Where full exploitation is a research project, ELTON bounds it: what is the attack worth, and can it reach you.
How much poisoned data it would take to shift the model toward an attacker-chosen outcome. ELTON anchors the fraction to published results, converts it to a record count, and checks that count against the write access the technical tests actually found: a reachable path, capacity, integrity controls, and audit logging.
The query budget a copycat model would need to reach useful fidelity, and whether your rate limits and output detail let it get there. If the API returns bare decisions without probabilities or logits, ELTON reports the estimate and does not attempt a surrogate build.
Whether training records can be reconstructed from model outputs, given the output format and model type. With white-box access, ELTON runs one scored gradient-ascent inversion against a single agreed target record. GAN-assisted inversion stays out of the base scope.
Working evasion inputs built with gradient methods (PGD, FGSM, C&W) where there is white-box access, or transfer and query methods where there is not. ELTON sweeps the perturbation budget across epsilon values to find the smallest change that flips the output, matched to the input modality.
A threshold attack on loss or confidence that separates records known to be in the training set from records known to be out, and reports the separation. It needs a labeled set of in-training and out-of-training samples, the train and holdout split.
A data feed shifted just under the detector’s thresholds, measuring how far performance can be walked off before an alert fires. If no drift detection exists, that missing control is the finding, and the dependent test does not run.
The feasibility work has hard edges, so the engagement stays priced and deliverable.
Each feasibility assessment lists the access it needs: model weights or gradients, training data, thresholds, labeled samples. You provide them, with working access, before testing starts, and ELTON confirms them at kickoff.
Where a prerequisite is missing or access does not work in the window, ELTON assesses the threat from the architecture, the technical findings and published results, records the gap as a scope limitation, and the activity is deemed delivered. The window does not slip and the fee does not change.
Where a control under test does not exist, no ingest monitoring, no drift detection, the missing control is the finding and the dependent test does not run. You are not charged to test something that is not there.
One set of evidence that travels across the frameworks arriving now.
An AI-Enabled Device Cybersecurity Assessment Report. Every test result and feasibility judgment tied to one of the seven named threats, with traceable evidence, ready to carry into the submission.
Findings carry CVSSv3 and CVSSv4 ratings where they apply, through the MDDT-qualified rubric. The rating is product-adjusted and evidenced, not a generic score.
Every finding runs the VEX-aligned lifecycle on the ELTON platform, alongside the rest of the product’s findings, and carries across releases.
Does FDA require AI-specific cybersecurity testing?
For a cyber device under section 524B, cybersecurity is already a requirement. FDA’s January 2025 draft guidance for AI-enabled device software functions then names seven AI-specific threats, in Section XII, and recommends that sponsors address them in cybersecurity risk management and testing. The guidance is still draft, so ELTON treats it as direction, not settled text, and ties the hard requirement to 524B.
What are the seven threats?
Data poisoning, model inversion or stealing, model evasion, data leakage, overfitting, model bias, and performance drift. Only model evasion happens on the deployed model at runtime. The other six begin in the training data, the model and inference layer, the pipeline, or postmarket monitoring.
Do you need access to our model and training data?
The hands-on technical testing does not. The feasibility assessments do, and each one lists exactly what it needs up front, model weights or gradients, training data, thresholds, labeled samples. If you cannot provide a prerequisite, ELTON assesses that threat from the architecture and the technical findings and records the gap, so the engagement still completes.
What is the difference between a technical test and a feasibility assessment?
A technical test is hands-on penetration testing that produces a reproducible finding with evidence. A feasibility assessment is a bounded evaluation of an advanced adversarial technique, one target, one batch, one feed, that answers whether the attack is practical against your architecture and rates the risk, without running a full research-grade attack.
Is this the same as the AI you use to find vulnerabilities?
No. The pipeline that finds vulnerabilities uses AI to build tools and decide which to run; your data never enters a prompt. AI-Enabled Device Testing is about the AI inside your device, and the seven ways it can be attacked. A pipeline, not a prompt, on our side; the model and its pipeline, on yours.
How does this map to our FDA submission?
The report maps every result to one of the seven named threats, with evidence, and findings carry CVSSv3 and CVSSv4 ratings through the MDDT-qualified rubric and a VEX-aligned status. A reviewer reads the AI-specific risks addressed, tested, rated and dispositioned, inside the 524B cybersecurity record.