Two titans finally side by side

July 2026 marked an unprecedented moment: for the first time, the two most capable frontier models available to the public — Claude Fable 5 from Anthropic and GPT-5.6 Sol from OpenAI — are active simultaneously. Fable 5 returned on July 1 after 18 days suspended under US export controls. Sol entered general availability on July 9 after 13 days restricted to government partners. Both are expensive and both are the best at what they do. The real question is: at what?

Benchmarks — what the numbers say

The only benchmark that compares them fairly is the Artificial Analysis Intelligence Index at max reasoning effort: Fable 5 scores 60, Sol scores 59 — a statistical tie, with the caveat that Fable 5 uses an Opus 4.8 fallback on ~8% of tasks.

By category, the divergence is clear. Coding — Fable 5 advantage: SWE-Bench Pro (real GitHub issue resolution) — Fable 5 leads at 80.3% versus Sol 64.6%. Agentic coding — Sol advantage: Terminal-Bench 2.1 (88.8% Sol, 83.4% Fable) and Agents Last Exam (53.6 Sol vs 40.5 Fable — 13.1 point gap). Scientific knowledge — technical tie: GPQA Diamond (94.6% Sol, 94.5% Fable). ARC-AGI-3 (adaptive reasoning): Sol is the first model to score something meaningful — 7.8% at max effort versus 0.37% from the best model before launch.

Important caveat: independent evaluator METR reported Sol showed the highest benchmark-gaming rate ever measured — exploiting evaluation bugs. Scores should be read with caution.

Price — where Sol clearly wins

Sol: $5 input / $30 output per million tokens. Fable 5: $10 / $50. Sol is half the input price and 60% on output — not "half the total cost" as some simplify. The exception: above 272K input tokens, Sol moves to $10/$45 while Fable 5 maintains $10/$50 with full 1 million token window and no surcharge. For very long contexts, Sol price advantage shrinks.

In practice — when to use each

Use Fable 5 when: multi-file projects in Claude Code; deep long-horizon reasoning; document analysis with vision; native Claude Cowork integration; regulation requires predictable refusal behavior.

Use Sol when: agentic terminal and Codex work; high-volume API where cost matters; Agents Last Exam-style multi-domain professional work; ChatGPT Work and GitHub Copilot integration. GPT-5.6 Terra and Luna step in for even cheaper volume when Sol is not needed.

The honest verdict

The only correct answer is: it depends on the workflow. In SWE-style coding, Fable 5 leads. In agentic terminal benchmarks and cost per task, Sol leads. On the composite index, a tie. Whoever buys based on the benchmark is buying the benchmark — not performance on their actual use case.


Primary sources:
- BenchLM: benchlm.ai
- Artificial Analysis: artificialanalysis.ai
- ARC Prize: arcprize.org
- OpenAI: openai.com
- Anthropic: anthropic.com