Skip to content
DevMeme
6789 of 7590
OpenAI Announces GPT-5.1-Codex-Max With Day-Long Autonomous Coding Sessions
AI ML Post #7445 · source on Telegram

OpenAI Announces GPT-5.1-Codex-Max With Day-Long Autonomous Coding Sessions

Description

A screenshot of a social media post from someone at OpenAI announcing GPT-5.1-Codex-Max, a model that 'can work autonomously for more than a day over millions of tokens.' The post states 'Pretraining hasn't hit a wall, and neither has test-time compute.' It congratulates teammates @kevinleestone and @mikegmalek. Below the text is a technical description of long-running tasks and a line chart comparing GPT-5.1-Codex and GPT-5.1-Codex-Max on a 'Verified (n=500)' benchmark. The X-axis shows 'Thinking tokens' from 0 to 35,000 and the Y-axis shows difficulty levels solved (low, medium, high, xhi). The Codex-Max model (darker line) significantly outperforms the regular Codex model at higher thinking token budgets, reaching 'xhi' difficulty at 35k tokens. The text describes how Codex-Max automatically compacts its session when approaching context window limits, pruning history while retaining the most important context

Comments

30
Anonymous ★ Top Pick GPT-5.1-Codex-Max can work for 24 hours straight on your codebase -- finally an engineer that matches the expectations in the job posting
  1. Anonymous ★ Top Pick

    GPT-5.1-Codex-Max can work for 24 hours straight on your codebase -- finally an engineer that matches the expectations in the job posting

  2. Anonymous

    An AI that runs for 24 hours? Perfect. It can finally finish running 'npm install' on our legacy frontend project

  3. dev_meme 9mo

    And also max is actually xhigh but please, move along, less questions

    1. @RiedleroD 9mo

      still an improvement I guess?

      1. dev_meme 9mo

        Meh You know that I'm LLMs lover But this looks like shareholder-pleasing bullshit / damage control Gemini-3 is innovation Release of Sonnet-4.5 and Opus 4/4.1 were innovation This max which is actually xhigh looks like total bullshit Needs testing though

        1. dev_meme 9mo

          The only sad part about gemini-3 is that I only have access to pro, and not to gemini-3-ultra 🌚🌚🌚 But I can live with that

        2. @RiedleroD 9mo

          less tokes for same result seems like an ok improvement, though obv the headline is bs

          1. dev_meme 9mo

            Here the problem - this graph is for n=500

            1. @RiedleroD 9mo

              meaning?

              1. dev_meme 9mo

                500 attempts

                1. @RiedleroD 9mo

                  oh, shitty sample size. alright sure

                  1. dev_meme 9mo

                    Its like averages success rates If you want better codegex DX you want to see zero-shot improvements

                  2. @RiedleroD 9mo

                    yknow, if this was a good graph, they'd show inaccuracy ranges on the graph…

                    1. dev_meme 9mo

                      Well, OAI is famous for their charts after all 😄

                      1. @RiedleroD 9mo

                        it's become a rather famous statement, but it's still true: never trust a statistic you haven't faked yourself

        3. @lord_nani 9mo

          Kinda missed it, why is gemini 3 considered innovation? Pretty strong word for a current state of the llms

          1. @chupasaurus 9mo

            Because funny numbers on stock exchanges go up

          2. @Algoinde 9mo

            innovation in finance unprecedented stock inflation

  4. dev_meme 9mo

    So sad about those remaining 5% engineers who debug mess left after 95%

    1. @RiedleroD 9mo

      I assume that includes "hey codex, refactor this- oh it ran rm -rf again"

  5. dev_meme 9mo

    Industry: PRs must be small and on point so its easy to properly review them OAI: IMAGINE HOW MUCH WORK IT CAN DO IN 24 HOURS OF CODEGEN?!

    1. dev_meme 8mo

      I come up with what useful it may work on for 24 hours It stopped after 10 minutes with 80% free context 😕

      1. @tema3210 8mo

        Tried 5.2?

  6. @RiedleroD 9mo

    but what for though

    1. dev_meme 9mo

      Yeah, thats the entire point of my rant in comments for this screenshot 😄

      1. @RiedleroD 9mo

        just wanted to put it into 4 words

    2. Bibi_28 9mo

      +1

  7. dev_meme 9mo

    Thats not so practically usable for codegen / doesnt look like something that will translate into meaningful improvement

  8. @qtsmolcat 9mo

    OpenAI hasn't yet learned more compute can't fix fundamental limitations. That adding 10k more GPUs doesn't make ChatGPT make shit up less

  9. @mrYakov 9mo

    Did they just add "try again" in system prompt and some sort of context summarisation ?

Use J and K for navigation