The benchmark built to resist
ARC-AGI-3 launched in March 2026 with an explicit mission: resist AI progress for as long as possible. Unlike its predecessors — ARC-AGI-1 (saturated by current models) and ARC-AGI-2 (where Sol scored 92%) — ARC-AGI-3 is an interactive benchmark: the model does not answer a static question but must explore an unknown environment, form hypotheses, test, and adapt in real time.
In March 2026, the best result on ARC-AGI-3 was 0.37%. On July 9, GPT-5.6 Sol scored 7.8% at max reasoning effort — becoming the first model to win a complete benchmark game (ft09, 87%). The average human scores 100%.
What Sol does differently
Per the ARC Prize Foundation, Sol does not score better because it executes better — it scores better because it orients better. "Sol can read an unfamiliar scene correctly and in the game own vocabulary. It treats a failed hypothesis as a reason to re-plan rather than thrash." Most agent failures happen upstream of the code — in the environment orientation phase. Sol breaks that pattern.
The cost is real: max effort consumed approximately $20,000 in evaluation tokens versus $10,000 for "high" effort. At low and medium efforts, Sol barely registers above zero. The score only rises meaningfully at xhigh and max — meaning the model differential behavior is compute-hungry and expensive.
What 7.8% actually means
The honest answer: a real advance, not a breakthrough. Benchmark creator François Chollet framed ARC-AGI-3 as a "multi-year project" — 7.8% is an opening move, not a finish line. What the number confirms: frontier models in 2026 are beginning to demonstrate situational adaptation in novel environments. What it does not confirm: general intelligence. The benchmark tests visual-logical adaptation in constructed environments — not language, long-term memory, causal reasoning, or cross-domain knowledge transfer.
The trajectory
ARC-AGI-1: saturated. ARC-AGI-2: Sol scored 92% at $1.44 per task. ARC-AGI-3: 7.8%. Each version was designed to be exponentially harder than the previous. Historical pattern suggests ARC-AGI-3 will also be saturated — the question is when and at what computational cost.
Primary sources:
- ARC Prize: arcprize.org
- Office Chai: officechai.com
- OpenAI: openai.com

