The Absurd Pace of AI Progress
Description
This image is a screenshot of a tweet from user Dan Hendrycks (@DanHendrycks). The tweet critiques a recent Apple paper that claims AI systems struggle with puzzles easy for humans, citing GPT-4o's 69.9% score versus humans' 92.7%. Hendrycks counters that the paper omitted recent reasoning models, stating that a model named 'o3' achieves 96.5%, surpassing human performance. The tweet includes a composite image: one part shows examples of cognitive tests from a research paper (like perspective taking and maze completion), and overlaid on the right is an unrelated, absurdist meme chart. The chart has a header 'JLO' and three rows with numbers '99.5', '50.5', and '92'. The technical humor lies in the rapid, almost daily obsolescence of major AI research findings. A significant paper from a tech giant like Apple is publicly refuted with new data, showcasing the breakneck speed of AI development. The 'JLO' chart is a non-sequitur, a form of niche internet humor or 'shitposting' common in tech circles, adding a layer of surrealism to the very serious topic of AI benchmarks
Comments
44Comment deleted
The half-life of an AI benchmark paper is now shorter than the time it takes to get CI to pass on a legacy codebase. By the time it's published, it's already a historical document
Impressive numbers, but I’ll believe o3’s reasoning supremacy when it can configure an Apple developer certificate on the first try
Apple: 'Our puzzles are so easy, humans ace them!' OpenAI: 'Hold my gradient descent.' Classic case of benchmarking against yesterday's models while the field moves at GPU clock speeds - by the time your paper's published, someone's already fine-tuned their way past your human baseline
Ah yes, the classic AI research cycle: publish a benchmark showing AI fails at 'simple human tasks,' then six months later watch reasoning models absolutely demolish it. It's like we're speedrunning Moravec's Paradox in reverse - turns out spatial reasoning puzzles are easier than getting an LLM to consistently count the letter 'r' in 'strawberry.' The real puzzle here is whether Apple's researchers intentionally skipped o3 to make a point, or if their paper review process moves at the speed of a waterfall SDLC. Either way, I'm sure the next benchmark will involve something truly human-specific, like understanding why we still use YAML despite its cursed indentation rules
Apple's benchmark rigor meets frontier model velocity: evals obsolete before arXiv hits production
99.5 on JLO, 50.5 on mazes - quintessential leaderboard engineering: P95 demo, P50 outage
Cool - o3 crushes JLO at 96.5%; wake me when a model can reason through a Friday rollback on a multi‑region microservice while finance asks why the AWS bill looks like a maze
I'm not following the situation, please explain Comment deleted
showel seller was shocked to find out that you can replace people with showels now Comment deleted
idiot finds out that by spending more non-renewable resources and time, you can replace humans with AI in certain tasks no one cares about, more at eleven Comment deleted
Calling LLM an AI is a disgrace to a human race Comment deleted
disrace Comment deleted
pls no racism Comment deleted
you started it. I'm just making some jokes Comment deleted
im a dog, i can be racist Comment deleted
Your website is racist towards mobile users. I can see only half of captcha and parts of a page are on top of each other 🙈 Comment deleted
guys, please do your dishes or something Comment deleted
The weak must suffer Comment deleted
A racing dog I assume? Comment deleted
no, backdoor labrador Comment deleted
That’s just heterophobic Comment deleted
I love racism. My favourite racist is Max Verstappen Comment deleted
SS (Super Super) Max Comment deleted
Can you play 5 stars on hidden mod? Comment deleted
what abut min verstappen Comment deleted
Apple got cucked out of mainstream AI market so now they're trying to gaslight everyone into thinking AI isn't actually that big of a deal Comment deleted
me when i don't know how science works Comment deleted
it aint lol Comment deleted
This Comment deleted
> big deal "oooo look at me I can generate unmaintainable code even faster now" Comment deleted
Apple was in consumer ML integration on their iphones before it was a thing. They did a lot of improvement that doesn’t look like ML did it but they are. Comment deleted
ofc why trying to improve apple ai, just spend time creating a paper on how ai isnt reasoning Comment deleted
I don't think people understand the Apple pullout well. They've always been about shipping well-polished and very detailed features that hold up to their standards, and not shipping something they think is subpar. If you remember, that was their whole thing about not having a calculator on iPad, because they couldn't just ship a boring calc, it had to be extra. They see that right now, AI cannot be relied on as a flawless tool that will hold up to their standards. Where they were able to apply ML to give consistent and almost deterministic results, they use it (like generative emojis). However, with LLMs they failed to achieve what they wanted (perfection), because LLMs are insanely imperfect at the moment for the amount of compute they reqiuire, so they decided it will be a net loss for them. That said, all pure conjecture. Comment deleted
I bought my iphone in 2018 and after airpod pro, i still used win and now i used mac. but apple need to improve on a lot of things Comment deleted
I also think that quite a bit of what they make is dogshit, but their intention is what I described. They just sometimes fail at it. Comment deleted
The refutation of the Apple's research: https://arxiv.org/html/2506.09250v1 Comment deleted
Copium Comment deleted
ask gippity summarize Comment deleted
Y'all kinda toxic Comment deleted
Here Comment deleted
Fox baby Comment deleted
~2kW/h for sudoku Bright future ahead Comment deleted
J/s² don't really makes sense as a unit, and neither does kW/h Comment deleted
JLO (no scores) They asked Jennifer Lopez to solve those puzzles, and she did complete none of them? Comment deleted