The same CVE in three different releases of your product can carry three different ratings, and all three can be correct. That is not a bug in the process. It is the point. A vulnerability's severity depends on the architecture, the controls, and the attack surface of the specific release it lands in.
This is why per-release context is the only honest way to score. A finding that is unreachable in the shipped v2.0 because a service was removed is genuinely not affected there, even if the same CVE is critical in v2.4 where the feature came back. Score the product as one blob and you will be wrong in both directions.
Every finding, rating, clock, and metric should be scoped to a release. When a finding moves between releases, it should carry forward as a clone with its own context, not a copy-paste of a rating that no longer applies. That is what makes postmarket metrics like time-to-remediation meaningful instead of misleading.
If your tool gives one product one score, it is telling you a comfortable fiction.
Start with one device. We build the twin from documentation your quality system already produces, run AI discovery remotely, and show you the graph: the handful to fix, and the evidence for everything else.