BullshitBench: Charting Which LLMs Actually Push Back on Users
Description
A line chart titled 'BullshitBench: How have models improved?' with subtitle 'Tracing performance improvements (clear pushback %) with model releases' and a checked option 'Best models from the same release'. The Y-axis is '% Clear Pushback (Green)' from 0% to 100%; the X-axis is 'Model Launch Date' from Q1 2024 to Q2 2026. Three vendor lines with dashed trend lines: Anthropic (orange) climbs from Claude 3 Haiku (~10%) through Claude 3.5 Sonnet (~45%), Claude Opus 4 / Opus 4.1 (dip to ~34%), Claude Sonnet 4.5 (~79%) up to Claude Sonnet 4.6 (~90%+); OpenAI (green) goes from GPT-4o Mini (~12%) through Gemini-adjacent mid-range to GPT-5.1 Chat, GPT-5.2 Codex, GPT-5.3 Chat and GPT-5.3 Codex hovering around 20-50%; Google (blue) tracks Gemma 3 27b IT (~3%), Gemini 2.5 Pro, Gemini 3 Pro, Gemini 3.1 Pro and Flash/Flash Lite variants mostly between 10-50%. The benchmark satirically measures sycophancy: how often a model clearly pushes back on a user's wrong premise instead of agreeing
Comments
14Comment deleted
Finally, a benchmark where 'You're absolutely right!' counts as a failing grade
The winning architecture may just be `assert premise.is_valid()` before generating 2,000 tokens.
100% means bullshit or push back? Comment deleted
Newer models are better at this as you can see Comment deleted
According to the Bench gpt 4-o mini have 0 resistance to bullshit. This bench is clearly sponsored by Claude 😅 Comment deleted
Tbh, from my use of the models claud is actually the best at correcting you and not just going along with stuff Comment deleted
Like what kind of stuff? Comment deleted
Claude is also very good in changing its mind multiple times in a single message Comment deleted
Love the "wait, but → actually → ⟲" loop depleting my token limits 15 minutes before the deadline ☺️ Comment deleted
AI learns out-of-order execution 👌 Comment deleted
I like AI. It's like a self-guided pistol that 20% hits the target, 79% misses it, and 1% hits your balls instead. But hey, it's so convenient and useful! 😀 Comment deleted
i wish that was that low Comment deleted
That google plot be like (source: https://xkcd.com/2048/) Comment deleted
Channels over 1000 followers get ads that fit the channel or something mostly scams i guess Comment deleted