Skip to content
DevMeme
7573 of 7590
Opus 5 Peaks at Medium, Then Overthinks
AI ML Post #8299 · source on Telegram

Opus 5 Peaks at Medium, Then Overthinks

Description

Two Cognition FrontierCode v1.1 line charts titled "FrontierCode (main set): test-time compute scaling" and "FrontierCode (extended set): test-time compute scaling." Legend: GPT-5.6 Sol (gray), Claude Opus 4.8 (blue), Claude Fable 5 (yellow), Claude Opus 5 (orange). X-axis is average cost per task in USD on a log scale (~$2–$20); Y-axis is score percent. Points are labeled low/med/high/xhigh/max. On the main set, orange Opus 5 jumps from low 41.9% to a peak at med 53.4%, then falls through high 48.0%, xhigh 43.6%, and max 48.0%. Yellow Fable 5 climbs to xhigh 53.5% then max 51.6%. Gray GPT-5.6 Sol tops out around max 47.5%; blue Opus 4.8 peaks near high 46.6%. The extended set repeats the shape: Opus 5 med 63.6% is the orange high-water mark before high 58.3%, xhigh 56.9%, max 58.9%. Captions note scores run by Cognition and that Opus 5's best scores are at medium effort. The meme is the non-monotonic curve: after med thinking, extra test-time compute buys worse coding results.

Comments

1
Anonymous ★ Top Pick Opus 5 at xhigh is just the model reopening a merged PR to add comments, then failing the tests it already passed at medium.
  1. Anonymous ★ Top Pick

    Opus 5 at xhigh is just the model reopening a merged PR to add comments, then failing the tests it already passed at medium.

Use J and K for navigation