Skip to content
DevMeme
6295 of 7590
AI ML Post #6901 · source on Telegram

The Absurd Pace of AI Progress

Description

This image is a screenshot of a tweet from user Dan Hendrycks (@DanHendrycks). The tweet critiques a recent Apple paper that claims AI systems struggle with puzzles easy for humans, citing GPT-4o's 69.9% score versus humans' 92.7%. Hendrycks counters that the paper omitted recent reasoning models, stating that a model named 'o3' achieves 96.5%, surpassing human performance. The tweet includes a composite image: one part shows examples of cognitive tests from a research paper (like perspective taking and maze completion), and overlaid on the right is an unrelated, absurdist meme chart. The chart has a header 'JLO' and three rows with numbers '99.5', '50.5', and '92'. The technical humor lies in the rapid, almost daily obsolescence of major AI research findings. A significant paper from a tech giant like Apple is publicly refuted with new data, showcasing the breakneck speed of AI development. The 'JLO' chart is a non-sequitur, a form of niche internet humor or 'shitposting' common in tech circles, adding a layer of surrealism to the very serious topic of AI benchmarks

Comments

44
Anonymous ★ Top Pick The half-life of an AI benchmark paper is now shorter than the time it takes to get CI to pass on a legacy codebase. By the time it's published, it's already a historical document
  1. Anonymous ★ Top Pick

    The half-life of an AI benchmark paper is now shorter than the time it takes to get CI to pass on a legacy codebase. By the time it's published, it's already a historical document

  2. Anonymous

    Impressive numbers, but I’ll believe o3’s reasoning supremacy when it can configure an Apple developer certificate on the first try

  3. Anonymous

    Apple: 'Our puzzles are so easy, humans ace them!' OpenAI: 'Hold my gradient descent.' Classic case of benchmarking against yesterday's models while the field moves at GPU clock speeds - by the time your paper's published, someone's already fine-tuned their way past your human baseline

  4. Anonymous

    Ah yes, the classic AI research cycle: publish a benchmark showing AI fails at 'simple human tasks,' then six months later watch reasoning models absolutely demolish it. It's like we're speedrunning Moravec's Paradox in reverse - turns out spatial reasoning puzzles are easier than getting an LLM to consistently count the letter 'r' in 'strawberry.' The real puzzle here is whether Apple's researchers intentionally skipped o3 to make a point, or if their paper review process moves at the speed of a waterfall SDLC. Either way, I'm sure the next benchmark will involve something truly human-specific, like understanding why we still use YAML despite its cursed indentation rules

  5. Anonymous

    Apple's benchmark rigor meets frontier model velocity: evals obsolete before arXiv hits production

  6. Anonymous

    99.5 on JLO, 50.5 on mazes - quintessential leaderboard engineering: P95 demo, P50 outage

  7. Anonymous

    Cool - o3 crushes JLO at 96.5%; wake me when a model can reason through a Friday rollback on a multi‑region microservice while finance asks why the AWS bill looks like a maze

  8. @exe0x0 1y

    I'm not following the situation, please explain

    1. @Turok1234 1y

      showel seller was shocked to find out that you can replace people with showels now

    2. @purplesyringa 1y

      idiot finds out that by spending more non-renewable resources and time, you can replace humans with AI in certain tasks no one cares about, more at eleven

    3. @Sp1cyP3pp3r 1y

      Calling LLM an AI is a disgrace to a human race

      1. @deadgnom32 1y

        disrace

        1. @Sp1cyP3pp3r 1y

          pls no racism

          1. @deadgnom32 1y

            you started it. I'm just making some jokes

            1. @Sp1cyP3pp3r 1y

              im a dog, i can be racist

              1. @Dark_Embrace 1y

                Your website is racist towards mobile users. I can see only half of captcha and parts of a page are on top of each other 🙈

                1. @purplesyringa 1y

                  guys, please do your dishes or something

                2. @Sp1cyP3pp3r 1y

                  The weak must suffer

              2. @kitbot256 1y

                A racing dog I assume?

                1. @Sp1cyP3pp3r 1y

                  no, backdoor labrador

                  1. @kitbot256 1y

                    That’s just heterophobic

          2. @Root16 1y

            I love racism. My favourite racist is Max Verstappen

            1. @anilakar 1y

              SS (Super Super) Max

              1. Sure Not 1y

                Can you play 5 stars on hidden mod?

            2. @RiedleroD 1y

              what abut min verstappen

  9. @OneAndOnlyMorgan 1y

    Apple got cucked out of mainstream AI market so now they're trying to gaslight everyone into thinking AI isn't actually that big of a deal

    1. @purplesyringa 1y

      me when i don't know how science works

    2. @itsTyrion 1y

      it aint lol

    3. @maks_mikh 1y

      This

    4. @itsTyrion 1y

      > big deal "oooo look at me I can generate unmaintainable code even faster now"

    5. @ZgGPuo8dZef58K6hxxGVj3Z2 1y

      Apple was in consumer ML integration on their iphones before it was a thing. They did a lot of improvement that doesn’t look like ML did it but they are.

  10. @ercolebellucci 1y

    ofc why trying to improve apple ai, just spend time creating a paper on how ai isnt reasoning

    1. @Algoinde 1y

      I don't think people understand the Apple pullout well. They've always been about shipping well-polished and very detailed features that hold up to their standards, and not shipping something they think is subpar. If you remember, that was their whole thing about not having a calculator on iPad, because they couldn't just ship a boring calc, it had to be extra. They see that right now, AI cannot be relied on as a flawless tool that will hold up to their standards. Where they were able to apply ML to give consistent and almost deterministic results, they use it (like generative emojis). However, with LLMs they failed to achieve what they wanted (perfection), because LLMs are insanely imperfect at the moment for the amount of compute they reqiuire, so they decided it will be a net loss for them. That said, all pure conjecture.

      1. @ercolebellucci 1y

        I bought my iphone in 2018 and after airpod pro, i still used win and now i used mac. but apple need to improve on a lot of things

        1. @Algoinde 1y

          I also think that quite a bit of what they make is dogshit, but their intention is what I described. They just sometimes fail at it.

  11. @Kumapawa 1y

    The refutation of the Apple's research: https://arxiv.org/html/2506.09250v1

  12. @TheFloofyFloof 1y

    Copium

  13. @Algoinde 1y

    ask gippity summarize

  14. @hur7m3 1y

    Y'all kinda toxic

  15. @hur7m3 1y

    Here

  16. @hur7m3 1y

    Fox baby

  17. @imfreetodowhatever 1y

    ~2kW/h for sudoku Bright future ahead

    1. _ 1y

      J/s² don't really makes sense as a unit, and neither does kW/h

  18. @SamsonovAnton 1y

    JLO (no scores) They asked Jennifer Lopez to solve those puzzles, and she did complete none of them?

Use J and K for navigation