Opus 5 Benchmark Card Versus Fable and Sol
Description
A clean comparison table of frontier AI models: Opus 5, Fable 5, Opus 4.8, and GPT-5.6 Sol. Opus 5 is framed in an orange box with many cells highlighted, showing claimed leads on Agentic terminal coding (Frontier-Bench v0.1, 43.3%), Knowledge work (GDPval-AA v2, 1861), Novel problem-solving (ARC-AGI-3, 30.2%), Agentic search (BrowseComp, 90.8%), Computer use (OSWorld 2.0, 70.6%), FrontierCode (53.4%), and AutomationBench (26.0%). Fable 5 edges Humanity's Last Exam no-tools (56.5%) and Legal (13.3%); GPT-5.6 Sol leads DeepSWE (72.7%). Health and Biology rows insert Mythos 5 scores under Fable. This is Anthropic's July 2026 Claude Opus 5 launch scorecard against Claude Fable 5, Opus 4.8, and OpenAI GPT-5.6 Sol.
Comments
1Comment deleted
43.3% looks like a landslide once you draw an orange box around the column and let Mythos 5 crash Fable's health row.