Skip to content
DevMeme
6434 of 7590
AI ML Post #7055 · source on Telegram

LLM Task Duration Chart Shows Exponential Growth in AI Agent Capabilities

Description

A scatter plot titled 'Task duration (for humans) where we predict the AI has a 50% chance of succeeding' with the Y-axis showing time from 0 to 2+ hours and X-axis showing LLM release dates from 2020 to 2026. Green dots represent specific models: GPT-3 (near 0), Claude 3.5 Sonnet Old (~15min), o1 (~35min), Claude 3.7 Sonnet (~50min), Claude 4 Sonnet (~1hr), o4-mini (~1h15m), o3 (~1h30m), Grok 4 (~1h45m), and GPT-5 (~2h15m). Gray dots with error bars show additional data points. A dashed green trendline shows exponential growth. Reference tasks on the Y-axis include 'Find fact on web' and 'Train classifier'

Comments

26
Anonymous ★ Top Pick At this rate, by 2028 AI agents will be able to handle tasks that take a human 8 hours - which means they'll finally be ready to attend a full day of meetings and still get nothing done
  1. Anonymous ★ Top Pick

    At this rate, by 2028 AI agents will be able to handle tasks that take a human 8 hours - which means they'll finally be ready to attend a full day of meetings and still get nothing done

  2. Anonymous

    The scary part isn't that GPT-5 can do a 2-hour task; it's that it will probably spend the first hour arguing about the requirements in a Slack channel, just like a human

  3. Anonymous

    If this curve holds, GPT-6 will finish your code review, draft the retro notes, and file a Jira ticket reminding you that humans are now the blocking dependency

  4. Anonymous

    Ah yes, the classic exponential curve where we're always just 2 years away from AGI - reminds me of how we've been 5 years away from fully autonomous vehicles since 2015. At least this chart is honest about the 50% success rate, unlike my junior devs claiming their code is 'production ready'

  5. Anonymous

    Ah yes, the classic 'AI will replace us by 2025' chart - where every new model release pushes the goalposts from 'find fact on web' to 'replicate two hours of human cognitive labor.' Notice how we've gone from GPT-3 barely managing web searches to GPT-5 allegedly matching a junior developer's afternoon sprint. The real joke? By the time GPT-5 actually ships, we'll have redefined 'human-level performance' to mean 'can survive a 4-hour architecture review meeting without suggesting we rewrite everything in Rust.' The dashed green line of exponential progress conveniently ignores that humans also get coffee breaks, context switching overhead, and existential dread - none of which scale logarithmically

  6. Anonymous

    By 2026 GPT-5 handles two-hour tasks at 50% success; the production pattern is idempotent retries with exponential backoff - until Finance pages us

  7. Anonymous

    Great, we finally automated two hours of work with a 50% SLO - marketing calls it AGI, SRE calls it a postmortem

  8. Anonymous

    Inference latency's true scaling law: O(parameters × existential_doubt)

  9. @apBUS_amp_K 1y

    Houston, the error bars have entered the space!!!

  10. @fkarg10 1y

    the task in question

    1. @Vlasoov 1y

      Generic answer more or less complicated Linux question

  11. @sysoevyarik 1y

    "50% chance of succeeding" OpenAI calculations is: it's always 50% chance. It succeeds or it doesn't.

    1. @ZgGPuo8dZef58K6hxxGVj3Z2 1y

      Bro they cant even scale graphs consistently

      1. @sysoevyarik 1y

        Oh I forgot 😁

  12. @sysoevyarik 1y

    Also OpenAI dudes too dumb to use logarithmic scale

    1. @Iizvullok 1y

      I think they chose to make it linear on purpose so it looks better.

    2. @deadgnom32 1y

      where? I see its quite linear here

    3. @deadgnom32 1y

      ah I get it. they could have used it, but they didn't. I think it's because most people can't properly comprehend log scale

  13. @Algoinde 1y

    And if you read the error bars, you will find the average LLM error rate

  14. @deerspangle 1y

    "Find fact on web" , 10 minute task. Ah yes, the contractor billing by time approach

    1. @SamsonovAnton 1y

      That's what my bosses think, especially on facts that nobody would publish on the web, at least for free. 🥲

  15. @rliskovenko 1y

    Procrastinating for two hours -- achievement unlocked

  16. @mrYakov 1y

    Its totally wrong comparing old models, that try to give answer instantly and modern agentic workflow, that consuming a lot of time rechecking everything. apply modern workflow to older models and their results will be much higher

  17. @jackietanyen 1y

    Error bars used to mean something

  18. @qtsmolcat 1y

    That 50% success rate is doing a lot of heavy lifting here

  19. @cacojo15 1y

    An example of task that takes 2h for a human: "Exploit a buffer-overflow in libiec61850" Source: https://metr.org/blog/2025-03-19-measuring-ai-ability-to-complete-long-tasks/

Use J and K for navigation