Recommended Free Tools
iTechGuides is reader-supported. When you buy through links on our site, we may earn an affiliate commission. As an Amazon Associate I earn from qualifying purchases. Learn more
Yes—at least for a narrow kind of visual puzzle, a model can produce an answer without writing out its intermediate reasoning in words. BDH-CQ uses recurrent memory and latent computation to tackle ARC-style tasks, and its authors report 29.5% pass@2 on the public ARC-AGI-1 evaluation set. That is evidence for this specific setup, not proof that AI reasoning generally happens without language or that the model is broadly intelligent.
What “reasoning without words” means in BDH-CQ
BDH-CQ is a model described in the preprint “BDH-CQ: In-Context Learning with Recurrent Latent Reasoning”, submitted to arXiv on August 10, 2026. It is designed to infer visual rules from examples, then apply a rule to a new puzzle.
At inference time, each example updates the model’s recurrent memory. When the query arrives, the model computes its answer iteratively in a high-dimensional latent space rather than generating a verbal chain of intermediate steps. In other words, “without words” refers to the model’s internal answer process in this task setup—not to an absence of input or output representations altogether.
What the ARC-AGI-1 result does—and does not—show
The authors report that their 150-million-parameter configuration reached 29.5% pass@2 on the public ARC-AGI-1 evaluation set. Pass@2 means the result allows two attempts; it should not be read as a one-shot success rate.
#1 Best Overall
ARC-AGI was introduced to probe how efficiently a system can acquire skills and generalize to unfamiliar tasks. The ARC Prize’s series description frames it around fluid intelligence, learning from limited experience, and unknown tasks. A score on ARC-AGI-1 is therefore relevant to rule learning from examples, but it is not a complete measure of intelligence, and it is not a result for every version in the ARC-AGI series.
BDH-CQ was built for ARC-style problems. Its reported score does not establish how it would perform across general-purpose reasoning, language, or everyday tasks, nor does it show that all AI models can reason without language.
What kinds of puzzles did it handle?
In its report on the work, Science News, published September 22, 2026, described stronger performance on some tasks involving turning or moving shapes than on tasks involving color changes or combinations of rules. It also reported that harder ordering and nesting puzzles improved after the model saw an example of similar difficulty.
These observations describe patterns in the tested tasks, not a universal ranking of puzzle types. They suggest that demonstrations can matter: seeing an example of comparable difficulty may help the model infer how to solve a more complex task.
Rank #3
Does avoiding written reasoning make AI cheaper?
The preprint authors report a computed inference cost of $0.0007 per task. This is their estimate for BDH-CQ under their stated setup, not a general price for AI reasoning. Science News noted that a comparison model’s cost was calculated differently, so the figures are not a like-for-like cost test; a direct savings ratio would be misleading.
Yuntian Deng, a computer scientist at the University of Waterloo, called the result “an interesting efficiency result” in comments reported by Science News. He also cautioned that it does not show BDH-CQ’s architecture is better than other approaches; more testing is needed to distinguish the effects of model design from training method. Jonas Geiping, a machine learning researcher at the ELLIS Institute Tübingen and the Max Planck Institute for Intelligent Systems, likewise noted that the model was made specifically for ARC-style problems, limiting direct comparison with general-purpose systems.
What is gained—and lost—when the model does not explain its steps?
A latent process can avoid generating a long textual chain of reasoning. That may be useful when the task is to compute an answer rather than communicate an explanation. But not producing intermediate text also makes it harder for a person to inspect how the model reached its answer.
Free tools Windows power users keep installed
One-click scans. No signup required.
There is a further caveat: written reasoning is not necessarily a faithful record of a model’s internal computation. Experts quoted by Science News emphasized both sides of the trade-off. As Deng put it, “Human language is useful for communicating reasoning, but it need not be the most efficient representation for every intermediate computation.” The absence of a verbal trace does not by itself prove either that the latent process is sound or that a text-generating model’s explanation reveals its actual internal steps.
Best Value
How strong is the evidence so far?
The work is an arXiv preprint, not a peer-reviewed paper in the status reported by Science News on September 22, 2026. The sources cited here do not establish later peer review or independent replication. The findings should therefore be treated as preliminary author-reported results, especially when drawing conclusions beyond the benchmark.
The best-supported conclusion is specific: BDH-CQ demonstrates that a model can solve some ARC-style visual reasoning tasks using recurrent latent computation without verbalizing intermediate reasoning. Whether that approach improves efficiency against alternatives under matched conditions, or transfers to broader kinds of reasoning, remains unsettled.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

