Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Yes—specialized AI systems have solved multiple International Mathematical Olympiad (IMO) problems. Google DeepMind reported that AlphaProof and AlphaGeometry 2 solved four of six IMO 2024 problems for 28/42 points, a silver-medal-equivalent score. In 2025, DeepMind said an advanced Gemini model with Deep Think reached gold-medal-standard performance using natural-language statements within the human 4.5-hour contest limit. Those results demonstrate major progress on fixed olympiad benchmarks, not a guarantee that AI can solve arbitrary unsolved mathematics.

What AI achieved at the IMO

IMO 2024: four of six problems and 28 points

DeepMind reported that its two-system combination solved four of the six IMO 2024 problems and earned 28 of the available 42 points. That total was equivalent to a silver medal under the IMO scoring rubric.

AlphaProof solved three of the five non-geometry problems: two algebra problems and one number-theory problem. Nature’s account identifies the number-theory result as the hardest problem in that year’s set. AlphaGeometry 2 solved the geometry problem.

AlphaGeometry 2 reportedly found its geometry solution in 19 seconds once it received a formalized version of the problem. DeepMind also reported that AlphaGeometry 2 solved 83% of geometry problems from the preceding 25 IMO years, a historical benchmark rather than an official contest result.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

IMO 2025: a stronger end-to-end claim

In a 2025 announcement, DeepMind said an advanced Gemini model with Deep Think achieved gold-medal-standard performance on IMO problems. The reported setup accepted the official problem descriptions in natural language and generated rigorous proofs directly within the 4.5-hour competition time limit.

This is a stronger comparison with a human contestant’s workflow than the 2024 setup, but it remains a company-reported evaluation on six fixed contest problems. It was not an officially administered human-competition entry or a medal awarded by the IMO jury.

How AlphaProof and AlphaGeometry 2 work

AlphaProof: reinforcement learning with Lean

AlphaProof is a reinforcement-learning system that trains itself to prove mathematical statements in Lean, a formal proof language and verification environment. A proof is accepted only when Lean’s kernel checks it, giving the system a machine-verified proof object rather than an plausibly worded argument.

Google Research describes test-time reinforcement learning in which the system generates and learns from many related problem variants during inference. This lets AlphaProof adapt its search to the structure of a particular problem instead of relying only on a fixed, pre-trained response.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

The 2024 AlphaProof solutions were formal Lean proofs. That formalization provides unusually strong checking, but it also means the natural-language olympiad statement had to be translated into the formal system before proof search could begin.

AlphaGeometry 2: language guidance plus symbolic geometry

Geometry requires different representations from algebra, number theory and combinatorics: diagrams, points, lines, angles and auxiliary constructions all interact. AlphaGeometry 2 combines language-model guidance with symbolic geometry reasoning and automatically generated auxiliary constructions.

For the IMO 2024 geometry problem, the system received a formalized statement before producing its solution. Its specialization explains why DeepMind treated geometry with a separate component rather than asking AlphaProof to handle every problem type.

Are the 2024 results comparable with human contestants?

The 28/42 total is meaningful because it uses the IMO scoring rubric, but the process was not identical to a student’s six-hour, two-day contest experience. The comparison depends on several separate dimensions:

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Dimension 2024 AlphaProof/AlphaGeometry 2 Human IMO contestant
Problem input Formalized statements were used for the reported solutions. Official statements are presented in natural language with diagrams or mathematical notation.
Proof form AlphaProof produced Lean-checked formal proofs; AlphaGeometry 2 used its geometry-solving pipeline. Contestants submit written mathematical arguments judged by human markers.
Specialization Separate systems handled non-geometry and geometry tasks. One person must handle every problem domain.
Time and compute The sources describe multi-day computational effort; total effort exceeded human contest-time constraints. Each contestant works under the official contest schedule and available personal resources.
Verification Lean checks formal proof steps, while the reported score was mapped to IMO points. Official IMO graders determine scores from submitted solutions.

DeepMind’s technical account says AlphaProof’s main training was halted and its hyperparameters were frozen before the official 2024 problems. That limits one form of adaptation, but it does not remove the distinction created by formalization, specialized components and computation that could continue beyond the human time window.

Why the 2025 Gemini result matters

The 2025 report changes two important conditions. Gemini with Deep Think reportedly worked from natural-language problem descriptions, and it produced proofs within the 4.5-hour limit used for the competition. Those conditions make the result closer to an end-to-end contestant workflow than the 2024 formalized, multi-day pipeline.

It still should not be described as an official AI IMO medal. The evidence is a reported performance on six known contest problems, not an independently administered competition with the same registration, supervision and judging procedures as human participants.

What these results prove—and what they do not

What is established

  • Specialized AI systems can solve a substantial fraction of very difficult olympiad problems.
  • Formal proof search can produce machine-checked solutions for algebra and number-theory tasks.
  • A dedicated geometry system can combine symbolic reasoning with learned guidance and auxiliary constructions.
  • At least one reported 2025 system operated from natural language under the stated 4.5-hour limit.

What remains unproven

  • These results do not show that AI can solve arbitrary unsolved problems in mathematics.
  • They do not establish independent mathematical research ability, such as choosing fruitful conjectures or developing a new theory over months or years.
  • A score on six fixed problems cannot establish universal theorem-proving reliability.
  • Success on a benchmark does not mean every generated proof is correct without the relevant checker or expert review.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

A practical way to compare AI olympiad claims

When a new result is announced, inspect the following six questions rather than relying on a medal label alone:

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
  1. Which domain? Separate geometry performance from algebra, number theory and combinatorics.
  2. What was the input? Check whether the system received the original natural-language statement, a formal specification, or additional annotations.
  3. What proof was produced? Distinguish a Lean-checked formal proof from a natural-language proof that still requires validation.
  4. What budget was allowed? Record wall-clock time, hardware, parallel searches and any computation performed before or after the nominal contest window.
  5. How many problems were solved? A total score, individual problem scores and partial-credit rules can tell different stories.
  6. Who verified it? Give greater weight to independently checked or officially administered evaluations than to an internal company report.

Does this mean AI can replace olympiad students or mathematicians?

No. Olympiad solving is a narrow but demanding benchmark. It rewards constructing a short argument for a known statement under strict scoring rules. Mathematical research also requires deciding which questions matter, finding definitions, recognizing promising patterns, communicating with collaborators and persisting through ideas that may fail repeatedly.

AI systems are already useful as proof-search assistants, formalization aids and sources of candidate lemmas. Their usefulness depends on the domain and on whether a human or proof checker can verify the output. The IMO results show that the frontier of machine mathematical reasoning has moved substantially; they do not remove the need for human judgment about correctness, relevance or originality.

The current bottom line

AI can solve genuine IMO-level problems, and the reported capability improved sharply between the 2024 and 2025 demonstrations. The 2024 AlphaProof and AlphaGeometry 2 result was four of six problems and 28/42 points, but it relied on formalization, specialized systems and computational effort beyond a human contest workflow. DeepMind’s 2025 Gemini result is closer to that workflow because it reportedly used natural language and the 4.5-hour limit. Both should be read as benchmark achievements with clearly defined conditions—not as evidence that AI has solved mathematics in general.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.