Skip to content
DevMeme
7150 of 7590
AI ML Post #7841 · source on Telegram

BullshitBench: Charting Which LLMs Actually Push Back on Users

Description

A line chart titled 'BullshitBench: How have models improved?' with subtitle 'Tracing performance improvements (clear pushback %) with model releases' and a checked option 'Best models from the same release'. The Y-axis is '% Clear Pushback (Green)' from 0% to 100%; the X-axis is 'Model Launch Date' from Q1 2024 to Q2 2026. Three vendor lines with dashed trend lines: Anthropic (orange) climbs from Claude 3 Haiku (~10%) through Claude 3.5 Sonnet (~45%), Claude Opus 4 / Opus 4.1 (dip to ~34%), Claude Sonnet 4.5 (~79%) up to Claude Sonnet 4.6 (~90%+); OpenAI (green) goes from GPT-4o Mini (~12%) through Gemini-adjacent mid-range to GPT-5.1 Chat, GPT-5.2 Codex, GPT-5.3 Chat and GPT-5.3 Codex hovering around 20-50%; Google (blue) tracks Gemma 3 27b IT (~3%), Gemini 2.5 Pro, Gemini 3 Pro, Gemini 3.1 Pro and Flash/Flash Lite variants mostly between 10-50%. The benchmark satirically measures sycophancy: how often a model clearly pushes back on a user's wrong premise instead of agreeing

Comments

14
Anonymous ★ Top Pick Finally, a benchmark where 'You're absolutely right!' counts as a failing grade
  1. Anonymous ★ Top Pick

    Finally, a benchmark where 'You're absolutely right!' counts as a failing grade

  2. Anonymous

    The winning architecture may just be `assert premise.is_valid()` before generating 2,000 tokens.

  3. dev_meme 5mo

    100% means bullshit or push back?

    1. @death_by_oom 5mo

      Newer models are better at this as you can see

      1. dev_meme 5mo

        According to the Bench gpt 4-o mini have 0 resistance to bullshit. This bench is clearly sponsored by Claude 😅

        1. @death_by_oom 5mo

          Tbh, from my use of the models claud is actually the best at correcting you and not just going along with stuff

          1. dev_meme 5mo

            Like what kind of stuff?

          2. @DerKnerd 5mo

            Claude is also very good in changing its mind multiple times in a single message

            1. @Iffycat 5mo

              Love the "wait, but → actually → ⟲" loop depleting my token limits 15 minutes before the deadline ☺️

  4. @SamsonovAnton 5mo

    AI learns out-of-order execution 👌

    1. @Ihor3056 5mo

      I like AI. It's like a self-guided pistol that 20% hits the target, 79% misses it, and 1% hits your balls instead. But hey, it's so convenient and useful! 😀

      1. @abra_mixabra 5mo

        i wish that was that low

  5. @TheUnknown007 5mo

    That google plot be like (source: https://xkcd.com/2048/)

  6. @SheepGod 5mo

    Channels over 1000 followers get ads that fit the channel or something mostly scams i guess

Use J and K for navigation