Skip to content
DevMeme
6392 of 7590
AI ML Post #7009 · source on Telegram

GPT-5 Performance Benchmark on Expert-Level Questions

Description

A bar chart with a white background, titled "Humanity's Last Exam (Full Set)* Expert-level questions across subjects". The y-axis is labeled "Accuracy, pass@1". The x-axis lists various AI models and configurations, such as "GPT-5 pro (python + search with blocklist)", "ChatGPT agent (browser + computer + terminal)", and "GPT-4o (no tools)". The bars are colored in shades of purple and pink to represent "With thinking" and white/light pink for "Without thinking". An OpenAI logo is visible in the top right corner. The chart shows GPT-5 pro with tools achieving the highest accuracy at 42.0%. This chart presents benchmark results from OpenAI, likely from a technical paper or presentation, evaluating the performance of their latest models, including the unreleased GPT-5. The data compares different model versions and their ability to answer expert-level questions, highlighting the significant performance boost when models have access to tools (like Python interpreters and web browsers) and "thinking" time (likely a reference to chain-of-thought or multi-step reasoning processes). For a senior developer, this is a direct, data-driven look at the progress of cutting-edge AI, relevant for understanding the capabilities and limitations of current and future large language models

Comments

7
Anonymous ★ Top Pick So 'thinking' is just another feature flag now? My entire career has been a beta test for this
  1. Anonymous ★ Top Pick

    So 'thinking' is just another feature flag now? My entire career has been a beta test for this

  2. Anonymous

    So it turns out the biggest model upgrade isn’t more parameters - it’s the checkbox labeled “actually think,” which, coincidentally, ships disabled in far too many prod deployments (human and silicon alike)

  3. Anonymous

    After 20 years of building distributed systems that barely hit 99.9% uptime, it's oddly comforting to see GPT-5 struggle to break 42% on expert questions - turns out even with unlimited compute, some problems remain stubbornly human, like explaining to stakeholders why the microservice mesh needs another refactor

  4. Anonymous

    Turns out 'thinking' is the killer feature we've been waiting for - GPT-5 with actual reasoning beats GPT-4o by 8x, proving that even AI needs to stop and think before answering. Though at 42% accuracy on expert questions, it's still performing at 'senior engineer confidently wrong in a design review' levels. The real plot twist? ChatGPT with a browser and terminal but no thinking outperforms OpenAI's o3 model with thinking - apparently sometimes it's better to just Google it and run the code than to philosophize about the solution

  5. Anonymous

    Even with 'thinking,' top LLMs score like juniors on a staff engineer oral exam - we're safe for another benchmark cycle

  6. Anonymous

    Apparently the best optimization isn’t quantization - it’s setting thinking_enabled=true and handing the agent Python+browser+search; amazing what a shell and a blocklist can do for pass@1

  7. Anonymous

    Pass@1 spikes the moment you enable thinking and hand it python+browser+terminal - so the frontier of AI remains a senior dev with a shell, a search bar, and permission to think

Use J and K for navigation