Skip to content
DevMeme
6393 of 7590
AI ML Post #7010 · source on Telegram

AI Model Benchmark: Humanity's Last Exam

Description

A dark-themed bar chart titled "Humanity's Last Exam (Full set)" presented on a slide with the main title "Humanity's Last Exam". The chart compares the performance of various AI models, with percentages on the y-axis. The models listed are "o3 (no tools)", "Gemini 2.5 Pro (no tools)", "Grok 4 (no tools)", "o3", "Gemini 2.5 Pro", "Grok 4", and "Grok 4 Heavy". The bars are colored grey and orange, with "Grok 4 Heavy" achieving the highest score of 44.4%. The "xI" logo, likely for xAI, is in the bottom left corner. This chart displays competitive benchmark results for large language models on a test set called "Humanity's Last Exam". It compares models from different AI labs, including Google (Gemini 2.5 Pro), xAI (Grok 4), and likely OpenAI ("o3"). The chart highlights the performance gains from using tools (as implied by the "no tools" labels for the lower-scoring models) and showcases "Grok 4 Heavy" as the top performer in this specific evaluation. For senior engineers, this is a snapshot of the ongoing "AI race," providing data points on how different models stack up against each other, which is crucial for making strategic decisions about technology adoption

Comments

7
Anonymous ★ Top Pick Another Tuesday, another 'Humanity's Last Exam' benchmark. At this rate, the real last exam will be figuring out which model's API has the fewest breaking changes this quarter
  1. Anonymous ★ Top Pick

    Another Tuesday, another 'Humanity's Last Exam' benchmark. At this rate, the real last exam will be figuring out which model's API has the fewest breaking changes this quarter

  2. Anonymous

    Sure, 44 % isn’t quite AGI - but it’s high enough that your CTO just asked if the annual performance review can be replaced with a curl call to /v1/chat/completions

  3. Anonymous

    Turns out 'Humanity's Last Exam' is just asking AI to correctly estimate a JIRA ticket - even Grok 4 Heavy with all its parameters can't crack the 50% mark on that impossible task

  4. Anonymous

    When your AI model needs tools to pass 'Humanity's Last Exam,' you realize we've successfully automated the art of looking up answers on Stack Overflow. The real question is: did Grok 4 Heavy achieve 44.4% by actually understanding the problems, or did it just get really good at parsing the documentation?

  5. Anonymous

    Top score 44.4% on “Humanity’s Last Exam” - apparently we’re grading on the press‑release curve, with tool‑use treated as ‘open book’ and confidence intervals left as an exercise for Legal

  6. Anonymous

    Gemini Pro with tools: -1% score. Even LLMs know bad integrations can introduce more bugs than they fix

  7. Anonymous

    AGI progress report: three orange bars, no error bars; statistically significant in Excel, production-ready in the board deck

Use J and K for navigation