A Skeptic's Guide to Modern LLM Research Papers
Description
This image is a screenshot of a tweet from the user 'darren' (@darrenangle). The tweet, set against a black background with white text, satirizes the perceived state of academic papers on Large Language Models (LLMs). The introductory line reads, 'LLM papers be like:'. Below this, it lists three fictional, humorously titled research findings. The first is 'ClearPrompt: Saying What You Mean Very Clearly Instead of Not Very Clearly Boosts Performance Up To 99%'. The second is 'TotallyLegitBench: Models Other Than Ours Perform Poorly At An Eval We Invented'. The third is 'LookAtData: We Looked At Our Data Before Training Our Model On It'. The humor critiques common tropes in AI research: presenting obvious conclusions as profound insights (prompt clarity), creating biased benchmarks that favor a specific model, and the critical issue of data contamination where models are inadvertently trained on their evaluation data. This resonates deeply with experienced engineers who are often critical of the 'publish-or-perish' culture and the hype cycle in AI/ML research
Comments
13Comment deleted
Some LLM papers read like they're just rediscovering the CAP theorem for natural language: you can have Clarity, Accuracy, or Performance, but good luck getting all three in a model that wasn't trained on its own test data
Some days it feels like the real breakthrough isn’t better models - it’s discovering ever more creative ways to cherry-pick a benchmark and slap an -GPT suffix on the title
After 20 years in the industry, I've learned that the length of an ML paper's title is inversely proportional to its actual novelty - and directly proportional to how many hyperparameters they tweaked until SOTA appeared
This perfectly captures the AI research paper industrial complex: where 'ClearPrompt' gets published for discovering that clarity improves results, 'TotallyLegitBench' somehow always shows your model winning on metrics you just invented, and 'LookAtData' reveals the shocking insight that examining your training data is actually useful. It's the academic equivalent of submitting a PR titled 'FixBug: Made the bug not happen anymore by fixing it' and expecting a promotion. The real innovation would be a paper called 'HonestEval: We Compared Against GPT-4 And Lost, But Here's Why Our Approach Still Matters' - but that one never makes it past peer review
If your ‘SOTA’ comes from a benchmark you wrote, with prompts you tuned after peeking at the test set, congratulations - you’ve reinvented p‑hacking with tokenizers
LLM papers' real innovation: crafting benchmarks so custom, only your model remembers the answers from 'pre-training'
If your benchmark ships the same corpus and tokenizer as your model, that’s not evaluation - it’s unit testing with the answers already in the fixtures
written by former kaggle users that were not bother themselves with calculation of validation metric Comment deleted
I've read a lot of research papers across various disciplines in last 6 years. Let's say LLM papers are 3/10 on "disrespect and disgrace" scale. When you get social sciences studies it's insulting even to bytes wasted on storage of said paper... Comment deleted
It's got some real problems with replication too. It's honestly surprising how bad the replication crisis is in computer science, (surely it should be trivial to replicate anything, right?), but then ML papers crank that up a notch Comment deleted
Oh, yeah... Putting reproduction as a criteria for papers then worryingly huge amount of papers wouldn't pass a smell test. What's a p-hacking and n=3 among friends where's there a paper to be published ;) Comment deleted
Those social science papers then used as justification to change lives of hundreds of millions of people, affecting even more 😭 Comment deleted
What’s the story behind LookAtData Comment deleted