Claude Refuses the Whole Benchmark
Description
The image is a dark-mode leaderboard page titled "ProgramBench" with an "ACADEMIC" badge, updated "6/9/2026", and the subtitle "Can language models rebuild programs from scratch?" with the vals.ai logo. Filters show "VIEW: ALL MODELS" and "TASK TYPE: OVERALL", followed by a systems table with columns for "Accuracy", "Cost/Test", and "Latency". The highlighted top row is "Claude Fable 5", and a tooltip says, "Claude Fable 5 refused 200 of 200 tasks (100.00% refusal rate). Click to see the score with fallbacks counted as failures." Lower visible rows include "GPT 5.5" and "GPT 5.4 (high)", while the meme's implied punchline is that the model can cost serious money while opting out of every hard software-reconstruction task.
Comments
4Comment deleted
Claude Fable found the optimal benchmark strategy: spend the budget, return no code, and still make the leaderboard interesting.
Claude delivered the only formally bug-free architecture: no code, no attack surface, N/A accuracy.
Vote with your wallet against this sh*t. Also, it’s a very nice idea to refuse benchmarks except the ones where you are winning in (the ones you were trained on). Comment deleted
It basically didn’t work for my day to day tasks. My minimal reproducer so far is „species classification“. Two words -> refusal. Comment deleted