What happened
The 2026 International Mathematical Olympiad concluded in the last week of July with a milestone that would have been unthinkable two years ago: three frontier language models scored a perfect 42/42 on the six problems of the world's most prestigious high school mathematics competition.
Deedy Das, partner at Menlo Ventures and an investor in Anthropic, built an independent evaluation harness and submitted the IMO 2026 problems to four frontier models. The results: Claude Fable 5 completed 42/42 in 1 attempt and just 1.8 hours — the fastest of the group. GPT-5.6 Sol required 1 additional attempt but proved the cheapest ($10–50 per run). Moonshot AI's Kimi K3 also hit 42/42, though it required 4 extra attempts and a significantly higher token volume. Axiom Math formally verified everything in Lean — the only system with machine-verifiable proofs.
The critical distinction: independent vs. official
The nuance most headlines missed: these scores were obtained via an independent harness, not through official IMO channels. Evaluation was handled by Claude agents rather than human judges who are former medalists or official IMO coordinators. Das's GitHub repository explicitly notes that scores should be viewed as strong indicators rather than authoritative results.
This contrasts with IMO 2025, where OpenAI and Google DeepMind had their solutions formally evaluated by IMO coordinators, both earning 35/42 — gold-medal tier performance.
The most important signal: saturation
The trajectory is unmistakable: silver in 2024, official gold in 2025 (35/42), and perfect independent scores in 2026. The IMO is becoming saturated as a benchmark for mathematical reasoning among frontier models. An analysis by vals.ai emphasizes that the IMO may no longer effectively differentiate top-tier models. Evaluators are shifting focus toward the International Olympiad in Informatics (IOI), which remains unsaturated. Claude Fable 5 currently leads the IOI benchmark with 72.25% accuracy.
What changes in practice
Olympiad-level mathematical reasoning is now accessible to any developer for $10–50 per full session via API. "Solves IMO problems" is no longer a unique marketing differentiator, as multiple models can achieve it. The evaluation landscape is rapidly pivoting toward harder, highly practical benchmarks aligned with real-world engineering and research tasks.

