AI Model Benchmark: Humanity's Last Exam
Description
A dark-themed bar chart titled "Humanity's Last Exam (Full set)" presented on a slide with the main title "Humanity's Last Exam". The chart compares the performance of various AI models, with percentages on the y-axis. The models listed are "o3 (no tools)", "Gemini 2.5 Pro (no tools)", "Grok 4 (no tools)", "o3", "Gemini 2.5 Pro", "Grok 4", and "Grok 4 Heavy". The bars are colored grey and orange, with "Grok 4 Heavy" achieving the highest score of 44.4%. The "xI" logo, likely for xAI, is in the bottom left corner. This chart displays competitive benchmark results for large language models on a test set called "Humanity's Last Exam". It compares models from different AI labs, including Google (Gemini 2.5 Pro), xAI (Grok 4), and likely OpenAI ("o3"). The chart highlights the performance gains from using tools (as implied by the "no tools" labels for the lower-scoring models) and showcases "Grok 4 Heavy" as the top performer in this specific evaluation. For senior engineers, this is a snapshot of the ongoing "AI race," providing data points on how different models stack up against each other, which is crucial for making strategic decisions about technology adoption
Comments
7Comment deleted
Another Tuesday, another 'Humanity's Last Exam' benchmark. At this rate, the real last exam will be figuring out which model's API has the fewest breaking changes this quarter
Sure, 44 % isn’t quite AGI - but it’s high enough that your CTO just asked if the annual performance review can be replaced with a curl call to /v1/chat/completions
Turns out 'Humanity's Last Exam' is just asking AI to correctly estimate a JIRA ticket - even Grok 4 Heavy with all its parameters can't crack the 50% mark on that impossible task
When your AI model needs tools to pass 'Humanity's Last Exam,' you realize we've successfully automated the art of looking up answers on Stack Overflow. The real question is: did Grok 4 Heavy achieve 44.4% by actually understanding the problems, or did it just get really good at parsing the documentation?
Top score 44.4% on “Humanity’s Last Exam” - apparently we’re grading on the press‑release curve, with tool‑use treated as ‘open book’ and confidence intervals left as an exercise for Legal
Gemini Pro with tools: -1% score. Even LLMs know bad integrations can introduce more bugs than they fix
AGI progress report: three orange bars, no error bars; statistically significant in Excel, production-ready in the board deck