Science

AI earns gold at math Olympiad but falls short of perfect human score

AI earns gold at math Olympiad but falls short of perfect human score

Introduction

In a landmark moment for artificial intelligence, experimental models from Google DeepMind and OpenAI have achieved gold-medal standard at the 66th International Mathematical Olympiad (IMO), held in Queensland, Australia, in July 2025. For the first time, AI systems solved five out of six notoriously difficult math problems under the official 4.5-hour time limit, earning 35 out of 42 possible points—placing them in the top tier of competitors. This performance signals a significant leap in AI reasoning, particularly in symbolic and abstract problem-solving. Yet, despite this achievement, the world’s most advanced AI could not match the perfect scores achieved by five teenage human competitors. Their flawless solutions underscore that while AI is catching up, fundamental gaps remain in genuine mathematical insight and creative logical construction.

Key Details

  • Google’s DeepMind and OpenAI independently developed AI systems that achieved gold-medal scores (35/42) at the 2025 IMO.
  • The AI models solved five of six Olympiad-level problems within the 4.5-hour time limit, a dramatic improvement from DeepMind’s 2024 silver-medal result (28/42).
  • Five human contestants achieved perfect scores of 42/42, demonstrating superior reasoning and error-free execution.
  • Both AI systems produced solutions that were rated as clear and logically coherent by official IMO graders.
  • The exact computational resources used remain undisclosed, raising questions about scalability and environmental impact.
  • The AI models were based on advanced reasoning architectures, including Google’s Gemini with ‘deep think’ capabilities and OpenAI’s next-generation reasoning framework.

Background

The International Mathematical Olympiad, founded in 1959, is the most prestigious math competition for high school students worldwide. Each year, over 600 of the brightest young mathematicians from more than 100 countries compete to solve six complex problems spanning algebra, combinatorics, geometry, and number theory. Problems are designed to require deep insight, elegant proof construction, and creativity—not just computational accuracy. Success at the IMO often predicts future leadership in mathematical research, with alumni including Fields Medalists like Terence Tao.

Until recently, AI struggled even with basic arithmetic verification, as demonstrated by well-documented cases where large language models like ChatGPT and Google Gemini failed simple multiplication tasks—e.g., incorrectly calculating 4596 × 4859. These errors stem from the probabilistic nature of AI language models, which generate answers based on token prediction rather than symbolic manipulation. In contrast, mathematical proofs demand exactness, logical consistency, and step-by-step justification—conditions poorly suited to standard AI architectures.

Impact Analysis

Despite these challenges, the 2025 results indicate a pivotal shift. According to Google’s DeepMind team, their updated model, an advanced version of Gemini with enhanced ‘deep think’ reasoning, was able to simulate multi-step logical deduction by generating and evaluating intermediate proof steps—akin to how humans ‘work through’ a problem. OpenAI’s model employed a similar iterative refinement process, allowing it to discard incorrect pathways and converge on valid proofs.

Gregor Dolinar, president of the IMO, remarked on the quality of AI-generated solutions:

“Their solutions were astonishing in many respects. IMO graders found them to be clear, precise and most of them easy to follow.”
This assessment is significant: clarity and rigor are hallmarks of mathematical excellence, and AI’s ability to produce them marks a departure from earlier, often incoherent or hallucinated responses.

Still, the fact that no AI achieved a perfect score highlights enduring limitations. The sixth problem—a combinatorial geometry challenge involving invariants and recursive structures—was solved by all five perfect scorers but eluded both AI systems. Experts suggest the issue may lie not in raw computational power, but in the AI’s inability to form intuitive leaps or recognize hidden symmetries without explicit training data.

Broader Context

The environmental cost of these AI breakthroughs cannot be ignored. While neither Google nor OpenAI disclosed the energy or water usage for their IMO runs, prior research indicates that training and running advanced AI models can consume as much electricity as small nations. A 2023 New York Times analysis estimated that large AI systems could soon use as much power as Argentina—or more. With tech giants expanding data center infrastructure, including plans to tap into fossil fuel reserves, the sustainability of AI progress is under scrutiny.

Moreover, the IMO test was administered in a controlled environment. Real-world mathematical research involves open-ended inquiry, failed attempts, and conceptual innovation—areas where AI still lags. As mathematician Timothy Gowers noted, “AI can verify a proof, but we have yet to see it invent a new field.”

Future Outlook

The race toward AI that can match or surpass human performance in abstract reasoning is accelerating. Google and OpenAI are likely to refine their models further, possibly achieving a perfect IMO score within the next few years. However, the ultimate benchmark may not be competition points, but the ability to generate novel theorems or solve long-standing open problems like the Riemann Hypothesis.

Meanwhile, educators and mathematicians are exploring how AI tools can assist rather than replace human thinkers. Hybrid workflows—where AI handles routine verification and humans provide strategic insight—could redefine how mathematics advances in the 21st century.

Conclusion

The 2025 IMO results represent both a triumph and a reality check for AI. Gold-level performance proves that machines are closing in on human-level problem-solving in structured domains. But the perfect scores of five teenagers remind us that creativity, intuition, and flawless logical execution remain uniquely human strengths—for now. As AI evolves, the challenge will not be to beat humans at math, but to deepen our understanding of intelligence itself.