The 99.9% figure most closely matching this claim is a statement by Chelsea and Westminster Hospital NHS Foundation Trust about Skin Analytics DERM in one NHS pathway: the Trust said it had “99.9% accuracy in ruling out melanoma.” It is not a general measure of how well AI detects cancer. The report does not provide enough detail about the metric or study to independently assess what that percentage means.
Where does the 99.9% figure come from?
In its 2024/25 annual report, Chelsea and Westminster Hospital NHS Foundation Trust said Skin Analytics DERM, launched at the Trust’s Chelsea site in December 2024, had “99.9% accuracy in ruling out melanoma.” The wording is specific to melanoma rule-out in that clinical pathway. It does not establish the performance of other AI systems, other cancers, or every patient use case.
The report also says the pathway discharges benign cases without dermatologist input and freed more than 35% of specialist appointments. Those are operational outcomes reported by the Trust, not estimates that can be applied to other services.
The public passage does not state the number or characteristics of patients assessed, the reference standard used to determine whether melanoma was present, a formal definition of “accuracy,” or a confidence interval. Without those details, the 99.9% figure cannot be interpreted as a verified sensitivity, specificity, or overall diagnostic accuracy rate.
What’s actually slowing this PC down?
Pick the symptom - the matching free tool is one click away.
#1 Best Overall
What does “99.9% accuracy” actually mean?
The word “accuracy” can conceal several different questions. A model may classify images correctly in a study, identify people who have a disease, help rule it out, or influence whether a patient is referred. Those are not interchangeable endpoints.
- Sensitivity: Among people who have the target condition, how often does the test return a positive result?
- Specificity: Among people who do not have the condition, how often does it return a negative result?
- Negative predictive value: Among people with a negative result, how many do not have the condition? This depends in part on how common the condition is in the tested population.
- Overall accuracy: The proportion of all evaluated cases classified correctly. Its meaning depends on the mix of positive and negative cases and on how the study selected them.
These measures require a clear target condition and a reliable reference standard—the method used to establish whether each case truly had the disease. The U.S. Food and Drug Administration’s 2007 guidance, Statistical Guidance on Reporting Results from Studies Evaluating Diagnostic Tests, recommends reporting paired measures such as sensitivity and specificity, with two-sided 95% confidence intervals and counts as well as percentages. It says: “FDA recommends you report measures of diagnostic accuracy (sensitivity and specificity pairs, positive and negative likelihood ratio pairs) or measures of agreement (percent positive agreement and percent negative agreement) and their two-sided 95 percent confidence intervals.”
For the DERM percentage, the Trust’s report does not provide the information needed to determine which of these interpretations applies. In particular, “99.9% accuracy in ruling out melanoma” should not be silently converted into a claim that the system catches 99.9% of melanomas.
Why AI cancer results cannot be compared as if they measured the same thing
Clinical AI tools are designed for different conditions, inputs, and decisions. A slide-reading aid, a cervical-screening method, a skin-lesion pathway, and software used during breast surgery do not answer the same clinical question. Their results also depend on the population studied, disease prevalence, clinical setting, reference standard, and whether the endpoint is an image classification, a referral decision, or a patient outcome.
| Example | Task and reported result | What the result does—and does not—show |
|---|---|---|
| Skin Analytics DERM, Chelsea and Westminster Hospital NHS Foundation Trust annual report 2024/25 | The Trust reported “99.9% accuracy in ruling out melanoma” in its pathway launched at the Chelsea site in December 2024. | The report passage does not give the study population, denominator, reference standard, formal metric definition, or confidence interval, so the percentage cannot be independently interpreted from that account alone. |
| Paige Prostate, FDA notice (2021) | The cited study found a 7.3% average improvement in cancer detection on individual slide images compared with pathologists’ unassisted reads. | The study did not evaluate the effect on final patient diagnosis. FDA described the software as an aid for qualified pathologists and noted false-positive and false-negative risks. |
| AI-based dual-stain cervical screening, NCI report (2020) | In samples from 4,253 people, the study reported better sensitivity and substantially higher specificity than Pap cytology; colposcopy referrals were about 42% versus about 60% for Pap. | This was a defined research comparison. The NCI said further regulatory approval would be needed before the fully automated method could be used to screen HPV-positive women. |
| Lumisight with an imaging system, FDA report (2024) | For helping locate residual cancer in the cavity after lumpectomy, FDA reported image-level sensitivity of 49.1% and specificity of 86.5%; 43% of patients had at least one false-positive image and 8% had at least one false-negative image. | These figures describe a specific intraoperative task and study. Image-level results are not the same as a patient-level diagnosis. |
The examples should not be ranked against one another by their percentages. They concern different cancers, clinical decisions, study designs, and endpoints. FDA guidance on computer-assisted detection in radiology likewise treats clinical performance assessment as specific to the device and task.
Does FDA authorization mean an AI system is error-free?
No. FDA says it regulates medical devices, including AI-enabled devices, based on intended use and technological characteristics. Its examples include imaging systems that provide diagnostic information for skin cancer. The agency reported that it had authorized more than 1,600 AI-enabled medical devices for marketing in the United States as of September 2026. That count covers many types of devices; it is not a count of cancer detectors and does not indicate a shared performance level.
Authorization or clearance applies to a particular product and intended use, not to AI cancer detection as a whole. FDA’s Paige Prostate notice, for example, describes software used as an aid for qualified pathologists rather than as a replacement for clinical judgment.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Is the 99.9% figure the same as a cancer test’s negative-result percentage?
No. A separate example illustrates why apparently similar percentages can describe different things. Cologuard Plus is a stool-based colorectal cancer screening test that looks for abnormal DNA and blood; it is not an AI cancer detector. Its manufacturer says the test is intended for adults aged 45 and older at average risk, reports that it detects 95% of colon cancers, and says that in a subset of 18,911 average-risk patients aged 45–84, a negative result had a 99.98% chance of meaning the person did not have colon cancer. That negative predictive value depends on disease prevalence, and a positive result should be followed up with colonoscopy.
A negative-result probability is not the proportion of all cancers detected. The Cologuard figures concern a different technology, condition, and use case, so they do not verify the DERM claim or establish an AI performance benchmark.
How to assess an AI cancer accuracy claim
Before treating a percentage as meaningful, look for the information that defines what was tested and how the result was measured. FDA guidance recommends describing intended use, the study population, target condition, reference standard, cases and exclusions, and results across relevant clinical sites and demographic or clinical subgroups.
- Identify the decision: Is the tool screening, triaging, detecting a finding on an image, assisting diagnosis, or predicting a patient outcome?
- Check who and what were tested: Look for the cancer type and stage, patient population, test setting, number of positive and negative cases, and any exclusions.
- Find the reference standard: Determine how researchers established the true cancer status. Testing against a benchmark that is not a reference standard may not support direct sensitivity and specificity claims.
- Look for paired results: Seek sensitivity and specificity, or another appropriate pair of measures, with confidence intervals and case counts—not a percentage without context.
- Check validation and endpoint: Ask whether results were independently tested at other sites and in relevant subgroups, and whether the study measured only image performance or a clinical outcome.
For the DERM figure specifically, the Trust’s published passage leaves those validation details unstated. The claim supports only the narrow wording the Trust used; it does not establish a universal rate for AI cancer detection.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.
The Tool Desk
Outbyte Driver Updater FREEScan for outdated or missing drivers - takes under a minuteDriver Scan →Outbyte PC Repair FREERepair Windows errors before they cause bigger problemsFix Now →

