Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Mammograms have traditionally been used to look for signs of breast cancer that may already be present. Researchers are also testing whether AI can analyze mammogram images to estimate a person’s chance of developing cancer later. That is a different job: a risk estimate is not a forecast of what will happen to one person, and it does not detect a future tumor.

Detection and future-risk prediction answer different questions

A mammography detection aid analyzes the current exam for suspicious findings that may indicate cancer is present now. A future-risk tool instead uses an image to estimate the probability or category of developing breast cancer over a period of time. The FDA describes this latter type as professional-use software with an intended use distinct from diagnosis or detection: it is not meant to diagnose, detect, treat, or guide interpretation of cancer.

That difference matters in practice. A risk estimate concerns likelihood across a future time horizon; it does not say that a tumor is visible today, nor can it identify with certainty who will develop cancer. A screening result and a future-risk estimate should therefore not be treated as interchangeable.

What the research can—and cannot—show so far

A systematic review published in 2026 examined 41 studies of mammography-based AI for future breast cancer risk prediction. The studies covered research published from January 1, 2012, through February 28, 2025, and all were retrospective. The review reported these median AUC values across studies:

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Prediction horizon Median AUC Source and qualification
Up to 2 years 0.71 Median across studies in the 2026 systematic review
3–4 years 0.72 Median across studies in the 2026 systematic review
5 or more years 0.71 Median across studies in the 2026 systematic review

AUC, or area under the receiver operating characteristic curve, describes how well a model distinguishes people who later develop cancer from those who do not. These medians summarize discrimination in the reviewed studies; they are not an individual’s probability of cancer, a guarantee of performance at a particular clinic, or evidence that using the tool improves health outcomes.

Why a risk score needs more than a strong AUC

Calibration: do the probabilities match real-world outcomes?

Discrimination and calibration are different. A model can rank people by relative risk yet still give probabilities that are too high or too low. Only six of the 41 studies in the 2026 review reported calibration, and reported results ranged from good calibration to overestimation of risk. That limited evidence makes it especially important not to read a model’s score as a dependable personal probability without knowing how it was calibrated for the population in which it is being used.

Representation: who was included in the evidence?

Most studies used 2D mammography images, and White, non-Hispanic women were the most represented group. Performance in a study population does not automatically carry over to people or screening facilities that differ from it. The review calls for more diverse populations, greater use of digital breast tomosynthesis (3D mammography), evaluation of aggressive or advanced cancers, and prospective studies.

External validation and clinical usefulness

A model should be tested on patients and facilities separate from those used to develop it, including groups that were less represented in earlier work. Researchers also need to establish whether risk estimates lead to better-informed screening or prevention decisions and, ultimately, meaningful patient outcomes. Retrospective performance results alone cannot answer those questions.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

What researchers are evaluating next

An NCI-funded project for fiscal year 2025 plans an independent evaluation of four commercial mammography-based risk algorithms at seven U.S. screening facilities. It is designed to examine performance across racial and ethnic groups and compare the algorithms with existing clinical risk-factor approaches. This is an evaluation plan, not proof that the algorithms improve care.

An NCI-listed active trial, “Artificial Intelligence Intervention for Improving Interpretation of Screening Mammography,” compares interpretation of 3D mammograms with and without AI and tracks immediate measures as well as one-year outcomes. It concerns ongoing evaluation of AI in screening interpretation; an active trial record is not a finding of improved patient outcomes.

The older evidence offers context but should not displace the more recent review. A 2024 Journal of the American College of Radiology review of 16 studies, whose literature search ran through September 30, 2022, reported a median AUC of 0.72 for image-only AI models versus 0.61 for density or clinical-risk-factor tools. In seven direct comparisons, six found no significant improvement when clinical factors were added to the image model. Those results reflect the earlier evidence base and do not establish which approach is best for a particular patient or current clinical setting.

How to judge a mammography-based risk model

When a study or clinical service describes an AI risk estimate, these questions help reveal what the result actually means:

Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
  • What is being predicted? Check the cancer type and time horizon, and whether cancers found on or near the mammogram used for prediction were included.
  • Where was the model tested? Look for external validation in populations and facilities distinct from the development data.
  • Who was represented? Ask whether performance was evaluated across racial and ethnic groups and other relevant populations, rather than inferred from an overall average.
  • Is the model calibrated? AUC measures discrimination, not whether the stated absolute risks match observed rates.
  • Which images were used? Evidence from 2D mammography does not automatically establish performance with 3D tomosynthesis.
  • Has it been tested in practice? Prospective evaluation should examine whether clinicians can use the estimate appropriately and whether decisions or patient outcomes improve.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

What a patient should do with an AI risk estimate

Ask the healthcare professional who ordered or discussed the assessment what population and time horizon the score applies to, how it was validated, and how it fits with other risk factors and screening guidance. NCI experts emphasize that risk estimates are population-based averages, not certain predictions for an individual. Ruth Pfeiffer, Ph.D., of the NCI, puts it plainly: “Unfortunately, these models cannot predict the future with certainty for any one individual.”

A high estimate does not mean cancer is inevitable, and a low estimate is not a reason to stop recommended screening. Decisions about screening schedules or preventive medication require discussion with a healthcare provider; medication also involves weighing possible side effects. As NCI expert Peter Kraft, Ph.D., says, “And it’s important to remember that these estimates of risk do not guarantee a specific outcome.”

What would make this a meaningful advance

The promise is that a mammogram might eventually contribute to a more individualized discussion of future risk, rather than serving only as an image to inspect for current signs of disease. Turning that possibility into useful care depends on independent validation, reliable calibration, representative evidence, and prospective tests of whether risk-informed choices help patients without worsening inequities. The studies described so far are part of that evaluation; they do not establish that AI-based future-risk assessment has improved outcomes or become a routine standard of care.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.