Do these 3 things before closing this tab:
1Repair Windows errors before they cause bigger problems2Fix the driver behind crashes, sound loss and screen glitches3Clear out junk files and repair common Windows errorsiTechGuides is reader-supported. When you buy through links on our site, we may earn an affiliate commission. As an Amazon Associate I earn from qualifying purchases. Learn more
AI systems can now solve some competition-level mathematics problems and produce formal proofs that software can check. But “AI solves math” covers several different achievements. The clearest published benchmark-style result comes from the 2024 International Mathematical Olympiad (IMO), a yearly competition for pre-university students. Google DeepMind’s AlphaProof and AlphaGeometry 2 solved four of six problems at a level equivalent to a silver medal. That is a real result on a defined task. It is not, by itself, evidence that AI can conduct broad mathematical research, explain why a result matters, or replace the expert reading that mathematicians rely on.
Whether a particular AI-generated proof is correct depends on how it was written. A formal proof can be checked mechanically by a proof assistant. An ordinary-language proof still needs expert reading.
Four tasks that all get called “solving math”
When a headline says an AI system “solved” a math problem, it may mean one of four different things. Each is judged differently, and success on one does not transfer automatically to another.
| Task | What the system produces | How correctness is judged | Coverage in the sources cited here |
|---|---|---|---|
| Answer production | A numerical or symbolic final answer | Compared with a known correct answer | IMO-Bench answer-accuracy dimension |
| Natural-language proof writing | A proof written in ordinary mathematical prose | Expert reading. The IMO-Bench project page states that human expert evaluation remains the gold standard for mathematical proofs. | IMO-Bench proof-writing dimension |
| Proof grading | A judgment about whether a given proof is correct | Assessed as a separate skill; the method is not described in the sources cited here | IMO-Bench grading dimension |
| Formal theorem proving | A proof written in a formal language such as Lean | A proof assistant accepts or rejects the formal artifact | IMO-Bench Lean proof dimension; 2024 IMO problems translated into Lean |
IMO-Bench, a benchmark published through the IMO-Bench project page, measures these dimensions separately rather than collapsing them into one score. Keeping them apart is the simplest way to avoid overstating what a system can do.
#1 Best Overall
The 2024 IMO result, with its conditions
The most cited milestone appears in the Artificial Intelligence Index Report 2025 from the Stanford Institute for Human-Centered AI. It reports that AlphaProof and AlphaGeometry 2 solved four of six 2024 IMO problems at a silver-medal-equivalent level. Three qualifications matter before you rely on it:
- The problems were manually translated into Lean. The systems therefore worked from formal versions that people prepared, not from an automatic conversion of the problem statements.
- The same report says it was then unknown how these systems would perform on traditional theorem-proving benchmarks. One olympiad result does not establish performance on those benchmarks.
- The conditions most relevant to a fair comparison, including time limits, tool use, and how much human steering was allowed, are not established by the source cited here. Treat the result as a result on its stated setup.
Where formal-math AI started
Formal-math AI predates the current wave of headlines. In a 2022 post, OpenAI described a Lean theorem prover solving selected high-school olympiad problems. That post is useful history: it shows that machine-checked proofs were already in use before 2024. It is not a current performance measure, so its results should not be compared with the 2024 or 2026 figures discussed here.
How Lean checks a proof
Lean is a proof assistant and programming language. You write a theorem statement and a proof in its formal language. Lean’s kernel then checks each step against the rules of the system and the definitions in the library. If the file compiles without errors, the proof is accepted under those rules. This is what “machine-checked” means.
What a passing check establishes
- The proof follows from the stated definitions, axioms, and inference rules.
- Every step has been checked by software rather than only read by a person.
What a passing check does not establish
- That the Lean statement says what the mathematician meant. A correct proof of a misformalized statement is still a correct proof of the wrong claim.
- That the result is significant, well motivated, or useful to other mathematicians. Verification is a separate question from importance. Whether a result is illuminating or a productive basis for further work is a judgment that people make.
- That the formal translation was faithful. In the 2024 IMO case, experts translated the problems by hand. Who wrote the formal statement is part of the evidence.
A four-step check for a Lean-verified claim
- Read the theorem statement first and compare it with the claim in plain language. Lean will accept a proof of a weaker or different statement without complaint.
- Search the file for
sorry. Lean accepts it as a placeholder and reports a warning that the declaration uses sorry, but nothing is proved. - Run
#print axiomson the main theorem to list the axioms it depends on. Commonly used core axioms such aspropext,Classical.choice, andQuot.soundare normal in Mathlib-based work. An unexpected custom axiom is a reason to stop and investigate. - Confirm the Lean and library versions the file targets, and check that it compiles in that environment. A file that compiles on one version may fail on another.
theorem add_swap (a b : Nat) : a + b = b + a := Nat.add_comm a b
#print axioms add_swap
theorem unfinished (a b : Nat) : a + b = b + a := by
sorry
The first theorem is a complete proof. The second compiles with a warning, which is exactly the case step 2 is designed to catch.
Five questions to ask about any AI math claim
- Which task was it: answer production, natural-language proof, proof grading, or formal proof?
- What conditions applied: time limits, tools, internet access, human steering, or extra attempts?
- Were the problems public, held out, or translated into a formal language by experts?
- How was correctness judged: an exact answer, expert review, or a proof assistant?
- Does the claim establish only correctness, or also explanation, novelty, and research value?
Benchmark conditions change quickly. A result without a date and an evaluation setup is hard to compare with anything else.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.What OpenAI says it is changing
In a September 21, 2026 announcement, OpenAI described an advisory group on mathematics and artificial intelligence. Its remit, as the company describes it, is to advise on the review and communication of emerging results and on academic and professional standards.
Rank #4
In an October 6, 2026 post, the company said it is sharing Lean formalizations of many of its proofs and consulting the independent advisory group about release practices. Both posts are the company’s own account of its process. They show that a governance step has been taken. They do not show that the wider mathematical community has accepted any specific result.
Free tools Windows power users keep installed
One-click scans. No signup required.
Quick Recap
What is still unsettled
- Which recent AI-generated research-level claims have been independently reviewed and accepted by the wider mathematical community. The sources cited here do not answer that for specific claims, so judge each claim on its own record.
- Whether results on olympiad problems carry over to research mathematics. Nothing cited here establishes that transfer, and it would need separate evidence.
- How working mathematicians’ practice will change. The sources show AI systems producing formal and informal proofs. They do not measure how mathematicians use these tools in their daily work.
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

