GPT-5 Coding Benchmarks: The Gap Widens
Description
A presentation slide featuring two distinct data visualizations comparing the performance of OpenAI's AI models. On the left, a line graph titled 'SWE-bench Verified' plots Accuracy against 'Average output tokens' for 'Real-world software engineering tasks.' It shows two ascending lines, one for GPT-5 and one for OpenAI o3, with GPT-5 consistently achieving higher accuracy at every token level, ranging from 'Minimal' to 'High' complexity. On the right, a bar chart titled 'Aider Polyglot' compares the 'Multi-language code editing' accuracy of GPT-5, OpenAI o3, and GPT-4.1. GPT-5 leads with 88% accuracy, followed by OpenAI o3 at 81%, and GPT-4.1 significantly behind at 52%. These charts are classic benchmark slides from a tech product launch, designed to quantitatively establish GPT-5's superiority in practical software development scenarios. For experienced engineers, this data is significant as it demonstrates advancements in concrete, valuable skills: solving real-world coding problems (SWE-bench) and sophisticated code manipulation (Aider Polyglot), signaling a new level of capability for AI-assisted development tools
Comments
7Comment deleted
These benchmarks are impressive, but I'm waiting for the 'Correctly Interprets Vague Jira Ticket' benchmark before I'm truly sold
If scaling to 14k tokens turns GPT-5 into a 78 % SWE-bench genius, I’m emailing finance to triple our AWS bill - clearly context windows are the new team of senior engineers
GPT-5 achieving 88% accuracy on polyglot code editing is impressive until you realize the remaining 12% is probably just missing semicolons in JavaScript that would've worked anyway
Ah yes, GPT-5 achieving 88% accuracy on code editing - finally, an AI that can refactor legacy code almost as well as the senior engineer who wrote it can remember why they made those architectural decisions in the first place. Though I notice it still needs 10k tokens to reach 'High' difficulty, which is roughly the same amount of context I need to explain to stakeholders why we can't just 'make it work like the demo.'
SWE-bench says accuracy climbs with output tokens - the only curve where verbosity beats minimalism; expect Finance to invent “prompt budgets” right after we approve 10k-token ADRs
More tokens boost accuracy like verbose PRs: occasionally fixes bugs, always hikes the bill
Apparently SWE-bench accuracy is just f(tokens); in prod the gradient gets clipped by timeout=30s and CFO_limit=medium