Claude model versions benchmarked on Pokémon Red milestones in staircase chart
Description
The image is a light-themed slide containing a line chart titled “Claude models playing Pokémon*” with the subtitle “Milestone progress over time.” The y-axis, labelled “MILESTONE REACHED,” lists (bottom to top) Start, Leave House, Get Starter, Reach Viridian City, Get Oak’s Parcel, Reach Viridian Forest, Get Brock’s Badge, Reach Mt. Moon, Reach Cerulean City, Get Misty’s Badge, Reach Vermilion City, and Get Surge’s Badge. The x-axis, labelled “NUMBER OF ACTIONS,” is marked from 0.0k to 35.0k in 5k increments. Four step-shaped lines compare model versions: a faint green “3.0 Sonnet” barely leaves the origin, a blue “3.5 Sonnet,” a purple “3.5 Sonnet (new),” and an orange “3.7 Sonnet” that climbs to the top milestone. A footnote cites Pokémon’s trademark, and a caption explains that Claude 3.7 Sonnet achieves the most game milestones with fewer interactions, framing the chart as an AI-performance benchmark for LLM-controlled gameplay
Comments
16Comment deleted
Claude 3.7 grabs Surge’s badge in fewer steps than our microservices need just to negotiate TLS - turns out RLHF beats three architecture review boards and a six-page ADR every time
Watching Claude grind through Viridian Forest is like watching a junior dev refactor legacy code - technically making progress, but you know there's a more efficient path if only they knew about Repel... or proper abstraction patterns
When your AI model's biggest achievement isn't passing the bar exam or acing medical boards, but finally getting past Brock without a water-type starter - truly the hardest benchmark in computer science. Claude 3.7 Sonnet: 35k actions to reach Surge's gym. My childhood self: 35k actions just trying to figure out how to leave Pallet Town
Counting 'actions' is basically story points for agents - Claude speedruns Kanto while our release process wipes on the first NPC
Impressive staircase, but until there's a fixed seed, variance bars, and an eval that rewards states instead of button‑mashing, it's basically a speedrun of the marketing deck
Claude 3.5 Sonnet scales Kanto gyms like params: linear progress, zero CAP theorem violations
I like to see the models play call of duty Comment deleted
They will never become racist enough unfortunately Comment deleted
AI competing humans in speedrun? Comment deleted
I find this chart particularly useless without human comparison Comment deleted
"the very best" I see what you did there. Comment deleted
I miss Twitch plays Pokemon :') Comment deleted
You mean AI needs 35k attempts for completing the game? How about Mario? Doom? Portal? Incredible Machines? Comment deleted
Not 35k runs, but 35k "interactions"; I presume this is the amount of button presses on the gameboy (emulator) Comment deleted
Starcraft gosus maximizing their apm rate must be mad about this. Comment deleted
Yeaap absolutely! Comment deleted