LLM Task Duration Chart Shows Exponential Growth in AI Agent Capabilities
Description
A scatter plot titled 'Task duration (for humans) where we predict the AI has a 50% chance of succeeding' with the Y-axis showing time from 0 to 2+ hours and X-axis showing LLM release dates from 2020 to 2026. Green dots represent specific models: GPT-3 (near 0), Claude 3.5 Sonnet Old (~15min), o1 (~35min), Claude 3.7 Sonnet (~50min), Claude 4 Sonnet (~1hr), o4-mini (~1h15m), o3 (~1h30m), Grok 4 (~1h45m), and GPT-5 (~2h15m). Gray dots with error bars show additional data points. A dashed green trendline shows exponential growth. Reference tasks on the Y-axis include 'Find fact on web' and 'Train classifier'
Comments
26Comment deleted
At this rate, by 2028 AI agents will be able to handle tasks that take a human 8 hours - which means they'll finally be ready to attend a full day of meetings and still get nothing done
The scary part isn't that GPT-5 can do a 2-hour task; it's that it will probably spend the first hour arguing about the requirements in a Slack channel, just like a human
If this curve holds, GPT-6 will finish your code review, draft the retro notes, and file a Jira ticket reminding you that humans are now the blocking dependency
Ah yes, the classic exponential curve where we're always just 2 years away from AGI - reminds me of how we've been 5 years away from fully autonomous vehicles since 2015. At least this chart is honest about the 50% success rate, unlike my junior devs claiming their code is 'production ready'
Ah yes, the classic 'AI will replace us by 2025' chart - where every new model release pushes the goalposts from 'find fact on web' to 'replicate two hours of human cognitive labor.' Notice how we've gone from GPT-3 barely managing web searches to GPT-5 allegedly matching a junior developer's afternoon sprint. The real joke? By the time GPT-5 actually ships, we'll have redefined 'human-level performance' to mean 'can survive a 4-hour architecture review meeting without suggesting we rewrite everything in Rust.' The dashed green line of exponential progress conveniently ignores that humans also get coffee breaks, context switching overhead, and existential dread - none of which scale logarithmically
By 2026 GPT-5 handles two-hour tasks at 50% success; the production pattern is idempotent retries with exponential backoff - until Finance pages us
Great, we finally automated two hours of work with a 50% SLO - marketing calls it AGI, SRE calls it a postmortem
Inference latency's true scaling law: O(parameters × existential_doubt)
Houston, the error bars have entered the space!!! Comment deleted
the task in question Comment deleted
Generic answer more or less complicated Linux question Comment deleted
"50% chance of succeeding" OpenAI calculations is: it's always 50% chance. It succeeds or it doesn't. Comment deleted
Bro they cant even scale graphs consistently Comment deleted
Oh I forgot 😁 Comment deleted
Also OpenAI dudes too dumb to use logarithmic scale Comment deleted
I think they chose to make it linear on purpose so it looks better. Comment deleted
where? I see its quite linear here Comment deleted
ah I get it. they could have used it, but they didn't. I think it's because most people can't properly comprehend log scale Comment deleted
And if you read the error bars, you will find the average LLM error rate Comment deleted
"Find fact on web" , 10 minute task. Ah yes, the contractor billing by time approach Comment deleted
That's what my bosses think, especially on facts that nobody would publish on the web, at least for free. 🥲 Comment deleted
Procrastinating for two hours -- achievement unlocked Comment deleted
Its totally wrong comparing old models, that try to give answer instantly and modern agentic workflow, that consuming a lot of time rechecking everything. apply modern workflow to older models and their results will be much higher Comment deleted
Error bars used to mean something Comment deleted
That 50% success rate is doing a lot of heavy lifting here Comment deleted
An example of task that takes 2h for a human: "Exploit a buffer-overflow in libiec61850" Source: https://metr.org/blog/2025-03-19-measuring-ai-ability-to-complete-long-tasks/ Comment deleted