Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

No: the 2018 study did not show that Google’s AI correctly predicted whether a patient would die 95% of the time. Its widely repeated “95%” was an AUROC score—a measure of how well a model ranked patients with an outcome above those without it across possible decision thresholds. The study reported strong retrospective discrimination for inpatient mortality at two US academic medical centers, not a 95% certainty about any one person or proof that using the model improved care.

What the study actually measured

Rajkomar and colleagues studied electronic health records for 216,221 adults who were hospitalized for at least 24 hours at two US academic medical centers. The models predicted several hospital outcomes, including whether a patient would die during the hospitalization. The work appeared in npj Digital Medicine in 2018.

For inpatient mortality prediction 24 hours after admission, the paper reported an AUROC of 0.95 (95% confidence interval 0.94–0.96) at Hospital A and 0.93 (95% confidence interval 0.92–0.94) at Hospital B. For comparison, the augmented Early Warning Score had AUROCs of 0.85 (95% confidence interval 0.81–0.89) and 0.86 (95% confidence interval 0.83–0.88), respectively. These figures compare models on the same outcome and prediction time within the study; they are not percentages of patients correctly classified.

Why AUROC is not “95% accuracy”

AUROC summarizes discrimination: how well a model tends to rank a randomly selected patient who experiences the outcome above a randomly selected patient who does not, considered across possible thresholds. A score of 0.95 does not mean that 95 out of 100 predictions were right. Nor does it mean that a patient given a high score has a 95% chance of dying.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

To turn a model score into a yes-or-no alert, someone must choose a threshold. Different thresholds change the balance between missed cases and false alarms. The share of predictions that are correct also depends on factors such as how common the outcome is in the population being assessed. AUROC alone does not tell a clinician which threshold to use or establish the practical consequences of acting on an alert.

The study assessed calibration separately, using predicted-versus-empirical probability curves. Calibration asks whether predicted risks correspond to observed outcome rates in groups; it is distinct from discrimination. Neither metric, by itself, establishes that the model is useful in clinical practice.

How the results were evaluated

This was a retrospective analysis of historical records, not a prospective trial of clinicians using the model. The researchers randomly assigned patients to development, validation, and test sets in an 80%/10%/10% split. The reported performance was measured on the held-out test set. That design evaluates performance on patients withheld from model development within the study data; it does not demonstrate that deploying the system changes care or patient outcomes.

What the paper does—and does not—support

Promising results in two study settings

The reported 24-hour AUROCs show that, on the study’s test sets, the models discriminated inpatient mortality better than the augmented Early Warning Score comparator. That is a meaningful research result, but it is specific to the hospitals, cohort, outcome, prediction horizon, and evaluation method used. Comparing the score with results from another hospital or study requires attention to those same factors, along with outcome prevalence and calibration.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

No demonstrated improvement in patient care

The authors cautioned against treating prediction quality as a clinical outcome: “Second, although it is widely believed that accurate predictions can be used to improve care, this is not a foregone conclusion.” A model may identify risk without showing what intervention should follow, whether a clinical team can act effectively, or whether the action helps patients.

Transfer to other hospitals was unresolved

Performance at two academic medical centers does not establish equivalent results elsewhere. The authors wrote: “Future research is needed to determine how models trained at one site can be best applied to another site.” Differences in records, patient populations, and workflows can matter when applying a model outside its development setting.

Not a consumer “death prediction” product

The study describes a research system built around electronic health records and computing infrastructure. Its authors said the FHIR-to-training pipeline and models depended on internal distributed computing platforms that could not reasonably be shared. The paper does not establish that this was a publicly available consumer product or an autonomous system issuing certain outcomes to patients.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

The accurate takeaway

“Google’s death AI” is a sensational shorthand for research that tested deep-learning models on hospital records. The paper found strong retrospective discrimination for inpatient mortality at two sites, including an AUROC of 0.95 at one hospital 24 hours after admission. It did not report 95% individual prediction accuracy, guarantee a patient’s outcome, or prove that deploying the model improves care.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.