Skip to content
DevMeme
6386 of 7590
AI ML Post #7000 · source on Telegram

GPT-5 Coding Benchmarks: The Gap Widens

Description

A presentation slide featuring two distinct data visualizations comparing the performance of OpenAI's AI models. On the left, a line graph titled 'SWE-bench Verified' plots Accuracy against 'Average output tokens' for 'Real-world software engineering tasks.' It shows two ascending lines, one for GPT-5 and one for OpenAI o3, with GPT-5 consistently achieving higher accuracy at every token level, ranging from 'Minimal' to 'High' complexity. On the right, a bar chart titled 'Aider Polyglot' compares the 'Multi-language code editing' accuracy of GPT-5, OpenAI o3, and GPT-4.1. GPT-5 leads with 88% accuracy, followed by OpenAI o3 at 81%, and GPT-4.1 significantly behind at 52%. These charts are classic benchmark slides from a tech product launch, designed to quantitatively establish GPT-5's superiority in practical software development scenarios. For experienced engineers, this data is significant as it demonstrates advancements in concrete, valuable skills: solving real-world coding problems (SWE-bench) and sophisticated code manipulation (Aider Polyglot), signaling a new level of capability for AI-assisted development tools

Comments

7
Anonymous ★ Top Pick These benchmarks are impressive, but I'm waiting for the 'Correctly Interprets Vague Jira Ticket' benchmark before I'm truly sold
  1. Anonymous ★ Top Pick

    These benchmarks are impressive, but I'm waiting for the 'Correctly Interprets Vague Jira Ticket' benchmark before I'm truly sold

  2. Anonymous

    If scaling to 14k tokens turns GPT-5 into a 78 % SWE-bench genius, I’m emailing finance to triple our AWS bill - clearly context windows are the new team of senior engineers

  3. Anonymous

    GPT-5 achieving 88% accuracy on polyglot code editing is impressive until you realize the remaining 12% is probably just missing semicolons in JavaScript that would've worked anyway

  4. Anonymous

    Ah yes, GPT-5 achieving 88% accuracy on code editing - finally, an AI that can refactor legacy code almost as well as the senior engineer who wrote it can remember why they made those architectural decisions in the first place. Though I notice it still needs 10k tokens to reach 'High' difficulty, which is roughly the same amount of context I need to explain to stakeholders why we can't just 'make it work like the demo.'

  5. Anonymous

    SWE-bench says accuracy climbs with output tokens - the only curve where verbosity beats minimalism; expect Finance to invent “prompt budgets” right after we approve 10k-token ADRs

  6. Anonymous

    More tokens boost accuracy like verbose PRs: occasionally fixes bugs, always hikes the bill

  7. Anonymous

    Apparently SWE-bench accuracy is just f(tokens); in prod the gradient gets clipped by timeout=30s and CFO_limit=medium

Use J and K for navigation