Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

iTechGuides is reader-supported. When you buy through links on our site, we may earn an affiliate commission. As an Amazon Associate I earn from qualifying purchases. Learn more

AI can flag suspicious claims, but it cannot reliably identify fake news in every topic, language, or breaking event. A model’s answer is not proof: trustworthy fact-checking also requires finding current evidence, judging its relevance, and explaining how it supports a conclusion. Research on benchmark classification, real-world news judgments, and AI-generated misinformation tests different things, so there is no single accuracy figure that tells you whether AI can detect fake news overall.

What does it mean for AI to detect fake news?

The phrase can describe several different jobs, and results for one should not be treated as evidence that a system can do all of them:

  • Classifying a headline or article: assigning a true/false or similar label, often against a benchmark with known answers.
  • Fact-checking a claim: identifying what is being asserted, finding relevant evidence, assessing whether that evidence supports it, and explaining the judgment.
  • Detecting AI authorship: estimating whether text was generated by AI. This is not the same as determining whether its claims are true; human-written content can be false, and AI-written content can be accurate.
  • Helping people judge news: improving readers’ ability to distinguish accurate from inaccurate stories. A system can label headlines well without improving readers’ decisions.

These distinctions matter because the studies below use different tasks, datasets, languages, and systems. Their results cannot be combined into a universal score for AI fake-news detection.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

What studies have found

Study and scope What it found What the finding does not establish
Hu et al., AAAI, published 24 March 2024: fake-news detection on two real-world datasets GPT-3.5 could often expose fake news and provide multi-perspective rationales, but underperformed a fine-tuned BERT model in the study. The authors’ ARG and distilled ARG-D methods outperformed three kinds of baseline on those datasets. A universal ranking of language models and specialist detectors. These are study-specific comparisons.
PNAS experiment: a specific ChatGPT version and a single prompt used to assess headlines and provide fact-checking information The tested LLM accurately identified most false headlines (90%) in that setup. Yet giving participants its information did not significantly improve their headline discernment or sharing of accurate news. Human-generated fact checks did improve discernment in the experiment. That current chatbots achieve 90% accuracy in real-world fact-checking, or that a good model label necessarily helps readers.
Suzgun et al., Nature Machine Intelligence, 2025: KaBLE, a benchmark of 13,000 questions across 13 epistemic tasks, evaluated with 24 language models The authors report systematic failures involving first-person false beliefs and weaker accuracy on those cases than on third-person false-belief cases. A direct, general fake-news accuracy score. The benchmark examines broader questions about belief, knowledge, and fact.
Ma et al., Nature Communications, online 11 December 2025; volume 17 (2026): Chinese-language datasets involving AI-generated deepfakes and cheapfakes The study reports limited intrinsic zero-shot detection capability in LLMs and finds that changes to linguistic features can cause detectors to fail. A numeric estimate of detection performance across all languages or every type of misinformation.

Why a correct-looking AI verdict can still mislead

In the PNAS experiment, the model’s fact-checking information did not significantly improve participants’ ability to discern accurate from inaccurate headlines. The researchers also identified harmful effects in particular cases: participants were less likely to believe true headlines the model mislabeled as false, and more likely to believe or share false headlines when the model expressed uncertainty about them. A label can therefore shape judgment even when it is wrong or inconclusive.

The authors describe fact-checking as a chain of tasks: detect claims, retrieve relevant evidence, assess whether the evidence establishes their veracity, and justify the conclusion. A fluent explanation is not a substitute for that chain. A model can produce plausible reasoning while misclassifying a claim, and a benchmark label alone does not show that readers will make better decisions.

Why breaking news is especially difficult

Developing stories create a freshness problem: a model may not have encountered the newest evidence, or may have learned an older account that has since changed. The PNAS authors call this the “breaking news problem”; they note that their model may have seen the study’s older false headlines but not newer true ones.

Access to trusted, current sources is a promising direction, the authors say, but their experiment does not show that retrieval solves the problem. A search result or linked source still has to be relevant, reliable, and interpreted correctly. Treat an answer about a live event as a lead to investigate, not a settled verdict.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

How AI detection compares with human news judgment

Human judgment is not a perfect benchmark either. A 2024 systematic review and meta-analysis in Nature Human Behaviour synthesized 67 publications, 195 samples, and 194,438 participants. It found average discernment between true and false news, alongside a smaller skepticism bias: participants were, on average, better at rejecting false news than affirming true news.

The review reported pooled discernment of d = 1.12 and skepticism bias of d = 0.32. These are standardized effect sizes, not percentages or AI performance scores. Its participant base was geographically uneven: 34% were from the United States, 54% from Europe, 6% from Asia, and 2% from Africa. Those proportions limit how confidently the findings can represent news judgment worldwide.

How to use AI when checking a story

Use an AI response to identify what needs checking, rather than as the final authority. The following steps apply the evidence requirements highlighted by the PNAS authors; their study did not test this checklist as a consumer method.

  1. Isolate the claim. Ask which specific factual assertion is being made, rather than whether an entire article is “fake.” Separate claims that can be checked from opinion, prediction, and interpretation.
  2. Open the underlying evidence. Do not rely only on the model’s summary or citations. Check that the linked material actually supports the claim and is not merely repeating it.
  3. Check source quality and context. Look for primary documents or independently corroborating reporting. Verify dates, location, quoted wording, and whether the evidence refers to the same event or people.
  4. Account for what may have changed. For a developing story, check whether the evidence is current. An answer that lacks current supporting sources should remain unverified.
  5. Treat uncertainty as a reason to investigate. A model’s confidence or uncertainty is not proof. If evidence is missing, conflicting, or irrelevant, do not convert the AI’s label into a confident conclusion.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

What to look for when comparing AI fact-checkers

Accuracy numbers are meaningful only in context. A useful comparison should test systems on the same dataset and language and distinguish headline or article classification from verification of individual claims. It should also report:

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
  • whether the system can access current external evidence and how well it retrieves relevant sources;
  • how often it falsely labels true claims and misses false ones;
  • whether its uncertainty is calibrated to its actual likelihood of being wrong; and
  • whether evaluation measures only model labels or also whether people make better judgments.

The studies discussed here are too heterogeneous for a head-to-head table of their accuracy figures. Their combined lesson is narrower but useful: AI can assist with detection and explanation in defined settings, while reliable fact-checking remains dependent on evidence, context, and evaluation of real reader outcomes.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.