Skip to content
DevMeme
6290 of 7590
AI ML Post #6894 · source on Telegram

AI Researchers Discover Models Can't Solve Unsolvable Problems

Description

The image is a screenshot of a fictional academic paper's abstract, formatted in a classic serif font on a white background. The paper is titled 'The Illusion of the Illusion of Thinking' and is presented as a comment on another paper by 'Shojaee et al. (2025)'. The abstract critiques the methodology of the original paper, which reported 'accuracy collapse' in Large Reasoning Models (LRMs). The core of the satire lies in the third point of the critique, which is highlighted in light blue: 'Most concerningly, their River Crossing benchmarks include mathematically impossible instances for N ≥ 6 due to insufficient boat capacity, yet models are scored as failures for not solving these unsolvable problems.' The humor is a dry, academic satire of the AI research field, specifically mocking poorly designed experiments and the rush to publish sensational findings. The punchline is the absurdity of blaming an AI model for failing to solve a problem that is logically and mathematically impossible, a scenario that resonates with engineers who often deal with flawed requirements or impossible-to-reproduce bug reports

Comments

21
Anonymous ★ Top Pick This is the academic equivalent of a JIRA ticket that says 'Bug: Application does not violate the laws of thermodynamics.' Priority: Blocker
  1. Anonymous ★ Top Pick

    This is the academic equivalent of a JIRA ticket that says 'Bug: Application does not violate the laws of thermodynamics.' Priority: Blocker

  2. Anonymous

    At last - a benchmark that mirrors our sprint planning: deliver the impossible by Friday or it’s a hard fail

  3. Anonymous

    When your paper claims LLMs can't reason but your experimental framework can't distinguish between 'model failed to solve' and 'you asked for a 10,000 token response with a 4,096 limit' - who's really exhibiting the illusion of thinking here?

  4. Anonymous

    When your benchmark suite includes mathematically impossible problems and you blame the AI for not solving them, you've successfully created a Turing test for detecting flawed evaluation frameworks. Turns out the real 'accuracy collapse' was in the experimental design all along - perhaps we need LRMs to peer-review our LRM benchmarks before we publish papers about LRM failures

  5. Anonymous

    Scoring LRMs as failures on N≥6 river crossings with a two‑seat boat is peak QA‑by‑Gödel - you’re not measuring reasoning, you’re benchmarking your impossible spec and the token budget

  6. Anonymous

    Scoring LLMs on N≥6 river-crossing is like calling Raft “broken” for not being linearizable during a partition - congrats, you benchmarked an impossibility theorem

  7. Anonymous

    LRMs master chit-chat but flop on Tower of Hanoi? Scaling laws meet their CAP theorem: you can't have reasoning, reliability, *and* reality

  8. bur del lago 1y

    a link would be interesting

    1. @offensive_otter 1y

      are you banned on google or something?

  9. @RiedleroD 1y

    I'm not sure what this is trying to say

    1. @deadgnom32 1y

      that researchers of AI sometimes spend so much time with AI — they forget how to think. they ask for solutions that can't exist and seeing the model failing to solve — mark it as a model's failure

      1. Sure Not 1y

        Just model this bruh.

      2. @RiedleroD 1y

        bruh

      3. @CcxCZ 1y

        And then some other researchers think asking for source code for a towers of hanoi puzzle, something that the model has to have at least thousand times in it's learning data, can be considered reasoning. Give me a break. Even computer security category on arxiv is spammed to hell with LLM-related junk. Also someone should add Towers of Hanoi to https://esolangs.org/wiki/HQ9%2B

        1. @deadgnom32 1y

          yes. I was also thinking about that.

  10. Sure Not 1y

    this.sense = new Someday()

    1. @CcxCZ 1y

      https://www.bootstrappable.org/

  11. @potompridumaiju 1y

    The paper implies that testers give LRMs the river crossing riddle with the number of elements that is more or equal to 6? Like, more than just a wolf, a sheep and a cabbage?

  12. @deadgnom32 1y

    basically — those researchers.

  13. @qtsmolcat 1y

    Waiting for the sequel, the illusion of the illusion of the illusion of thinking

    1. @qtsmolcat 1y

      But to be serious, LLMs are really big and really fancy text predictors (that cost fortunes to run). Of course it can't "actually" think, it's only simulating it

Use J and K for navigation