Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

In OpenAI’s 2024 SimpleQA test of short factual questions, the evaluated models often failed to provide a correct answer—but the results are not a universal error rate for AI. OpenAI’s later system-card table reports SimpleQA accuracy from 0.07 for o1-mini to 0.47 for o1, alongside different hallucination rates. Those figures describe particular model versions on a narrow benchmark, not every response from today’s AI systems.

What SimpleQA tested

OpenAI introduced SimpleQA on October 30, 2024, as an open benchmark for factuality. It contains 4,326 short, fact-seeking questions designed to have one verifiable answer. Topics include science and technology, television, and video games. OpenAI designed the questions to challenge frontier models while keeping evaluation relatively straightforward.

Trainers researched the questions and answers. A second trainer independently answered each question, and the benchmark retained questions for which their answers matched. A third trainer checked a random sample of 1,000 questions. After reviewing disagreements, OpenAI estimated that approximately 3% of the dataset had an inherent error. That is the authors’ estimate of possible errors in the benchmark itself—not a model’s error rate.

How the benchmark counted answers

SimpleQA distinguishes three outcomes, which matter when interpreting a score:

Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
  • Correct: The response gives the reference answer.
  • Incorrect: The response contradicts the reference answer. Hedging does not change this classification.
  • Not attempted: The response omits the reference answer without contradicting it.

That last category is not the same as a wrong answer. A system that abstains more may have fewer incorrect responses but also fewer correct ones. Accuracy, hallucination rate, and willingness to abstain describe different aspects of behavior; no single figure captures overall reliability.

What OpenAI reported for the tested models

OpenAI’s October 2024 SimpleQA publication evaluated GPT-4o-mini, o1-mini, GPT-4o, and o1-preview without browsing. It reported that the smaller models answered fewer questions correctly than GPT-4o and o1-preview, while the o-series models more often returned “not attempted.” OpenAI’s December 5, 2024 o1 system card later published a comparison table for five named models:

Model version in OpenAI’s table SimpleQA accuracy SimpleQA hallucination rate
GPT-4o 0.38 0.61
o1 0.47 0.44
o1-preview 0.42 0.44
GPT-4o-mini 0.09 0.90
o1-mini 0.07 0.60

These are the results OpenAI reported for the specific versions and benchmark setup in its system card, not a current ranking of model families. The table’s accuracy and hallucination columns are separate measures, so they should not be treated as complementary percentages that necessarily add to 100. The different values for GPT-4o and o1-mini, for example, illustrate why one metric alone cannot describe how a model handles every question.

Futurism’s November 2, 2024 article reported o1-preview’s SimpleQA success rate as 42.7%. OpenAI’s later system-card table gives that model’s accuracy as 0.42. These are source-specific representations of the benchmark result, not a timeless estimate of how often o1-preview—or AI generally—gets things wrong.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Can a model tell when it does not know?

SimpleQA’s “not attempted” label captures whether a response withholds an answer rather than contradicting the reference. OpenAI reported that the o-series models abstained more often than the other models in its original comparison. Abstention can be useful when a system lacks a reliable answer, but it is distinct from getting an answer right.

OpenAI also examined confidence. In its reported analysis, confidence and accuracy were positively related, but the models’ confidence remained short of perfect calibration: on average, they overstated how often their answers were correct. Confidence can therefore offer some signal without being a guarantee. The finding does not mean every confident answer is false or that confidence is never informative.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

What the results do—and do not—say about AI reliability

SimpleQA tests short questions with a single verifiable answer. It does not directly measure reliability when a model writes a long response containing many factual claims, handles specialist work, answers questions about changing facts, uses web browsing, or works across every subject area. OpenAI itself says whether performance on short factual answers correlates with the ability to write lengthy, fact-filled responses remains an open research question.

Best Value
Sale
Hands-On Machine Learning with Scikit-Learn, Keras, and TensorFlow: Concepts, Tools, and Techniques to Build Intelligent Systems
  • Use scikit-learn to track an example ML project end to end
  • Explore several models, including support vector machines, decision trees, random forests, and ensemble methods
  • Exploit unsupervised learning techniques such as dimensionality reduction, clustering, and anomaly detection
  • Dive into neural net architectures, including convolutional nets, recurrent nets, generative adversarial networks, autoencoders, diffusion models, and transformers
  • Use TensorFlow and Keras to build and train neural nets for computer vision, natural language processing, generative models, and deep reinforcement learning
  • Use the scores as benchmark evidence: They show how particular named model versions performed under this test’s conditions.
  • Do not turn them into a universal error rate: The figures do not establish how often any AI system is wrong across all tasks or in ordinary use.
  • Keep the measures distinct: Correct answers, incorrect answers, abstentions, hallucination rates, and calibration each address a different question.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.