Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Yes—early users reported that OpenAI’s o1-preview made conspicuous mistakes on tasks that looked simple, including counting the R’s in “strawberry,” solving a river-crossing puzzle, and making legal chess moves. But these examples, published in September 2024, were anecdotes rather than a controlled test, so they do not establish how often the model made such errors or how current models perform.

“Strawberry” was the reported codename; OpenAI called the public release o1-preview. Its launch benchmarks showed strong performance on selected math and coding evaluations, but those scores measured different things from the user-reported puzzles.

What mistakes did users report?

A September 13, 2024 report by Victor Tangermann at Futurism collected early-user examples of o1-preview appearing to fail at basic-seeming tasks:

  • Letter counting: users said the model struggled to count the letter R in “strawberry.”
  • River-crossing puzzle: Meta AI scientist Colin Fraser reportedly said the model abandoned a correct answer while working through the puzzle’s constraints.
  • Chess: INSA Rennes researcher Mathieu Acher attributed illegal moves to the model.
  • Logic puzzle: users reported getting varying answers to a strawberry-themed puzzle.

The article also included a user-reported 92-second response to a riddle and a “75 percent” result tied to one prompt. Neither figure is a representative latency measure or an estimate of the model’s overall accuracy.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Do the anecdotes prove o1-preview was broadly unreliable?

No. They show that some users observed errors in specific interactions. The examples were not a controlled, representative evaluation, and the cited reporting does not provide an independent error-rate estimate for letter counting, chess, or puzzle solving. A handful of striking failures cannot tell readers how frequently the model made those mistakes across prompts and users.

The converse is also true: strong benchmark scores do not prove that a model will answer every ordinary question correctly. The user reports and the benchmarks address different tasks and come from different kinds of evidence.

What did OpenAI’s launch benchmarks measure?

When OpenAI introduced o1-preview on September 12, 2024, it reported results on selected math and coding evaluations. These are company-reported scores for the specified tests, not general accuracy rates.

Launch-era result What OpenAI reported What it does not establish
IMO qualifying exam OpenAI reported an 83% score for its o1 reasoning model and 13% for GPT-4o on the same evaluation. It is not an error rate for chess, letter counting, puzzles, or everyday use.
Codeforces competitions OpenAI reported coding performance at the 89th percentile. It is not a measure of general factual reliability or puzzle-solving accuracy.

These figures appeared in OpenAI’s September 12, 2024 launch announcement. They support the narrower claim that o1 performed strongly on those evaluations; they do not resolve how reliably it handled the tasks in the user reports.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

What did OpenAI say about o1-preview at launch?

OpenAI presented o1-preview as an early model trained to spend more time thinking, refine its process, try strategies, and recognize mistakes. It also warned that the preview lacked some ChatGPT features, including web browsing and file or image uploads. OpenAI said, “For many common cases GPT‑4o will be more capable in the near term.” That was the company’s description of the products at launch in September 2024, not a statement about their capabilities today.

Do the 2024 reports describe current o1 models?

Not by themselves. The reports concern early use of o1-preview around its September 2024 release. OpenAI’s o1 system card, updated December 5, 2024, explains that its evaluations covered specified checkpoints and that production performance can vary with system updates, final parameters, system prompts, and other factors. The card describes the o1 family as using large-scale reinforcement learning to reason with chain-of-thought.

The system card’s preparedness ratings concern categories such as persuasion, chemical and biological risks, cybersecurity, and model autonomy. They are safety assessments, not scores for everyday factual accuracy or the specific simple-task errors in the 2024 anecdotes. Without a comparable, representative evaluation of later versions on those same prompts, the early reports should not be treated as evidence of their current error rate.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

How should readers interpret the “idiotic mistakes” claim?

The phrase captures the surprise of seeing an advanced model apparently stumble over a simple-looking prompt. The evidence supports a narrower conclusion: early users reported striking failures in particular o1-preview interactions, while OpenAI reported strong results on selected math and coding benchmarks. Neither set of evidence measures the model’s overall reliability, and neither makes the other disappear.

Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.