AI Researchers Discover Models Can't Solve Unsolvable Problems
Description
The image is a screenshot of a fictional academic paper's abstract, formatted in a classic serif font on a white background. The paper is titled 'The Illusion of the Illusion of Thinking' and is presented as a comment on another paper by 'Shojaee et al. (2025)'. The abstract critiques the methodology of the original paper, which reported 'accuracy collapse' in Large Reasoning Models (LRMs). The core of the satire lies in the third point of the critique, which is highlighted in light blue: 'Most concerningly, their River Crossing benchmarks include mathematically impossible instances for N ≥ 6 due to insufficient boat capacity, yet models are scored as failures for not solving these unsolvable problems.' The humor is a dry, academic satire of the AI research field, specifically mocking poorly designed experiments and the rush to publish sensational findings. The punchline is the absurdity of blaming an AI model for failing to solve a problem that is logically and mathematically impossible, a scenario that resonates with engineers who often deal with flawed requirements or impossible-to-reproduce bug reports
Comments
21Comment deleted
This is the academic equivalent of a JIRA ticket that says 'Bug: Application does not violate the laws of thermodynamics.' Priority: Blocker
At last - a benchmark that mirrors our sprint planning: deliver the impossible by Friday or it’s a hard fail
When your paper claims LLMs can't reason but your experimental framework can't distinguish between 'model failed to solve' and 'you asked for a 10,000 token response with a 4,096 limit' - who's really exhibiting the illusion of thinking here?
When your benchmark suite includes mathematically impossible problems and you blame the AI for not solving them, you've successfully created a Turing test for detecting flawed evaluation frameworks. Turns out the real 'accuracy collapse' was in the experimental design all along - perhaps we need LRMs to peer-review our LRM benchmarks before we publish papers about LRM failures
Scoring LRMs as failures on N≥6 river crossings with a two‑seat boat is peak QA‑by‑Gödel - you’re not measuring reasoning, you’re benchmarking your impossible spec and the token budget
Scoring LLMs on N≥6 river-crossing is like calling Raft “broken” for not being linearizable during a partition - congrats, you benchmarked an impossibility theorem
LRMs master chit-chat but flop on Tower of Hanoi? Scaling laws meet their CAP theorem: you can't have reasoning, reliability, *and* reality
a link would be interesting Comment deleted
are you banned on google or something? Comment deleted
I'm not sure what this is trying to say Comment deleted
that researchers of AI sometimes spend so much time with AI — they forget how to think. they ask for solutions that can't exist and seeing the model failing to solve — mark it as a model's failure Comment deleted
Just model this bruh. Comment deleted
bruh Comment deleted
And then some other researchers think asking for source code for a towers of hanoi puzzle, something that the model has to have at least thousand times in it's learning data, can be considered reasoning. Give me a break. Even computer security category on arxiv is spammed to hell with LLM-related junk. Also someone should add Towers of Hanoi to https://esolangs.org/wiki/HQ9%2B Comment deleted
yes. I was also thinking about that. Comment deleted
this.sense = new Someday() Comment deleted
https://www.bootstrappable.org/ Comment deleted
The paper implies that testers give LRMs the river crossing riddle with the number of elements that is more or equal to 6? Like, more than just a wolf, a sheep and a cabbage? Comment deleted
basically — those researchers. Comment deleted
Waiting for the sequel, the illusion of the illusion of the illusion of thinking Comment deleted
But to be serious, LLMs are really big and really fancy text predictors (that cost fortunes to run). Of course it can't "actually" think, it's only simulating it Comment deleted