Skip to content
DevMeme
5886 of 7590
AI ML Post #6445 · source on Telegram

AI Coding Assistant Benchmark: A Horse Race of Models

Description

A bar chart from the Aider code editing benchmark, displaying the performance of various large language models. The vertical axis is labeled 'Percent completed correctly' and ranges from 0 to 90. The horizontal axis lists several AI models. The results show 'claude-3-5 sonnet-20241022' as the top performer with approximately 84% accuracy. It is followed by 'o1-preview' at around 79%, 'DeepSeek V2.5' at about 72%, and 'gemini-exp-1206' highlighted in a distinct red bar at roughly 69%. Two other Gemini models, 'gemini-exp-1114' and 'gemini-exp-1121', score lower, around 61% and 58% respectively. The technical context, as provided by the original post's caption, is that this chart represents a significant benchmark for AI coding capabilities. It reflects the ongoing 'model wars' among major AI players like Anthropic (Claude), OpenAI (o1), and Google (Gemini). For experienced engineers, this isn't just a chart; it's a snapshot of a rapidly evolving tool-space where today's leader could be tomorrow's legacy system

Comments

12
Anonymous ★ Top Pick This benchmark shows which AI is best at editing code. The real-world benchmark is which AI is best at deciphering a three-year-old Jira ticket titled 'Fix the thing'
  1. Anonymous ★ Top Pick

    This benchmark shows which AI is best at editing code. The real-world benchmark is which AI is best at deciphering a three-year-old Jira ticket titled 'Fix the thing'

  2. Anonymous

    Nothing like a pastel bar chart to remind leadership that our entire Q3 roadmap depends on whichever model clears 70 % in a synthetic eval none of us volunteered to maintain

  3. Anonymous

    Gemini-exp-1206 in pink because it's embarrassed about being the only model update that made things worse - the classic 'we fixed the bug that was accidentally making it work' deployment

  4. Anonymous

    When your experimental model's version number goes up but the accuracy goes down - it's like deploying to production on a Friday, except the rollback takes three months and costs millions in GPU hours. The gemini-exp-1206 sitting there in pink is basically the architectural decision you defended in the design review that everyone's now too polite to mention in retros

  5. Anonymous

    Amazing how a 3% delta with no sample size or confidence intervals can trigger a platform-wide model swap, three ADRs, and a weekend of prompt‑router whack‑a‑mole

  6. Anonymous

    All these bars are probably inside the confidence interval we didn’t plot - yet procurement will mandate the pink one for ‘strategic alignment.’

  7. Anonymous

    Gemini-exp: Optimized for exponential hype, minimal completion

  8. @qtsmolcat 1y

    Claude's character limits kinda limit its coding usefulness tbh

  9. @SamsonovAnton 1y

    Aider Code editing benchmark Aides for everyone!

  10. @imfreetodowhatever 1y

    Why are devs the only ones so passionate with making themselves obsolete 🤔

    1. @qtsmolcat 1y

      That would require clients to a) know what they want and b) be able to clearly articulate that

  11. @SSS_Krut 1y

    Not agree, o1 made better results for me

Use J and K for navigation