iTechGuides is reader-supported. When you buy through links on our site, we may earn an affiliate commission. As an Amazon Associate I earn from qualifying purchases. Learn more
A model can infer a pattern from examples or derive a conclusion from stated premises. The harder question is whether it can invent a new premise that explains what it observes—the creative leap Tom Zahavy argues current large language models lack. His 2026 position paper makes a case, not a settled finding or proof that AI discovery is impossible.
Induction, deduction, and abduction: three different moves
These terms describe different ways of reaching a conclusion. A simple illustration—not an experiment—shows why they matter:
- Induction: You see several metal objects expand when heated and infer that heating may cause metal to expand. The conclusion generalizes beyond the examples, so it is plausible rather than guaranteed by them.
- Deduction: Given the premise “all metals expand when heated” and the fact “this object is metal,” you can validly conclude that this object expands when heated. The reasoning depends on the premises; deduction alone does not establish that they accurately describe the world.
- Abduction: You observe an unexpected result and propose a new explanation that could account for it. The proposal may be useful, but it still needs testing.
In Zahavy’s framing, the distinction is not that models can never infer or reason. It is that induction and deduction do not, by themselves, explain how a system originates a new explanatory premise.
Outdated Drivers Are Slowing You Down
One free scan finds every outdated or missing driver and matches the right update for your exact hardware.Free scan · exact hardware matchWindows Errors? Fix Them Before They Spread
Repair common Windows errors and clear accumulated junk for a smoother, more stable PC - no reinstall needed.Free scan · no reinstallWhat Zahavy means by “the jump”
In his 2026 ICML position paper, “Position: LLMs can’t jump”, Tom Zahavy frames scientific discovery as a move from experience or simulation to new axioms or explanatory hypotheses, followed by deductive reasoning from those premises. He uses Einstein’s formulation of general relativity as a computational case study and identifies translating simulation into formal axioms as a bottleneck.
#1 Best Overall
The paper’s strong claim—that current LLMs lack the mechanism for this abductive move—is Zahavy’s argument, not an established consensus or a demonstrated impossibility theorem. He writes: “We identify the translation of simulation into formal axioms as the critical bottleneck in artificial scientific invention, and propose that physically consistent, multimodal world models offer the necessary sensory grounding to bridge this divide.” Such world models are a proposed direction, not a proven solution.
What empirical studies say about generalization and rule-making
Task-specific studies complicate any broad claim that models cannot generalize. They examine particular tasks, methods, and evaluation settings; they do not directly settle whether a model can originate a scientific explanatory framework.
Rank #2
Symbolic tasks can expose limits as patterns grow
Jing Qian and colleagues’ 2023 ACL paper, “Limitations of Language Models in Arithmetic and Symbolic Induction,” reports that performance on certain copy, reverse, and addition tasks falls as the number of symbols or repetitions increases. Methods tested before the authors’ tutor-based approach did not completely solve their simplest addition induction problem. Their tutor-based method reached 100% accuracy in the paper’s tested out-of-domain and repeating-symbol situations. That figure applies to those specified settings and method, not to language models generally.
Recommended Free Tools
Generating a rule is not the same as applying it reliably
In “Phenomenal Yet Puzzling: Testing Inductive Reasoning Capabilities of Language Models with Hypothesis Refinement,” Linlu Qiu and colleagues examine iterative hypothesis refinement. Models propose candidate rules, a symbolic interpreter tests them against examples, and the models refine their proposals. The authors report that this hybrid process can produce strong results on several benchmarks, while also finding brittleness and gaps between generating rules and applying them.
Training can improve transfer in a tested setting
A 2026 preprint by Mingzi Cao and colleagues, “Fundamental Reasoning Paradigms Induce Out-of-Domain Generalization in Language Models,” reports that training on reasoning trajectories improved performance on its realistic out-of-domain tasks, with gains of up to 14.60 in that evaluation. The result shows transfer gains under the paper’s setup; it does not establish unrestricted generalization or settle the question of scientific invention.
Why these findings do not settle whether AI can make discoveries
“Can a model generalize?” has no single answer without specifying the task and conditions. Learning a pattern from examples, deriving a result from premises, proposing a candidate rule, and inventing a scientific explanation are related but distinct abilities. Evidence for one does not automatically demonstrate the others.
- What is being measured? A symbolic sequence task, a formal derivation, or the creation of an explanatory hypothesis?
- How novel is the test? Does it use familiar examples, out-of-domain cases, or sparse observations?
- What help is available? Is the model working alone or with a tutor, symbolic interpreter, or other tool?
- What kind of claim is being made? A philosophical thesis, an observed result on a particular benchmark, or a proposal for future research?
A 2022 TMLR paper’s Google Research record on emergent abilities adds a separate caution: some abilities are near random at smaller scales and cannot be predicted by extrapolating a scaling law from those models. This is a warning about forecasting capabilities from small-scale measurements, not proof of an unbounded or specifically abductive ability.
Do these 3 things before closing this tab:
1Scan for outdated or missing drivers - takes under a minute2Repair Windows errors before they cause bigger problems3Fix the driver behind crashes, sound loss and screen glitchesWhat remains open
Producing a novel hypothesis is not the same as validating it against observations, and validation is not by itself proof that the hypothesis is a useful scientific explanation. The cited studies offer evidence about particular forms of symbolic induction, rule refinement, and out-of-domain transfer. Zahavy’s paper raises the broader question of where new explanatory premises come from. The available results do not establish either that current systems can make scientific discoveries in this stronger sense or that they are incapable of doing so.
Quick Recap
Best Value
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

