Skip to content
DevMeme
6387 of 7590
AI ML Post #7002 · source on Telegram

GPT-5's Tool Use: Not a Universal Upgrade

Description

A bar chart from a presentation, titled 'T²-bench: tool use', comparing the accuracy of three AI models - GPT-5, OpenAI o3, and GPT-4.1 - across three different industry domains: Telecom, Retail, and Airline. The y-axis represents accuracy in percent. In the Telecom category, GPT-5 shows a massive performance lead with 97% accuracy, far ahead of o3 (58%) and GPT-4.1 (34%). In Retail, the models are much closer, with GPT-5 at 81%, o3 at 80%, and GPT-4.1 at 74%. In a surprising twist, for the Airline category, OpenAI o3 slightly outperforms GPT-5 with 65% accuracy compared to GPT-5's 63%. This chart provides a nuanced view of AI model performance, specifically on their ability to use tools. It demonstrates that while GPT-5 offers a monumental improvement in some areas (Telecom), its superiority is not absolute across all domains. For a senior technical audience, this highlights the critical importance of domain-specific testing and validation, as a newer model might not be the best choice for every single task, especially when specialized tool integration is involved

Comments

7
Anonymous ★ Top Pick The Airline benchmark proves what we've known for years: no amount of intelligence can make sense of a legacy GDS API from the 1980s
  1. Anonymous ★ Top Pick

    The Airline benchmark proves what we've known for years: no amount of intelligence can make sense of a legacy GDS API from the 1980s

  2. Anonymous

    Funny how the slide brags about 97 % accuracy in Telecom but stays silent on the 300 % increase in your GPU bill - classic ‘we’ll fix it in ops’ benchmarking

  3. Anonymous

    GPT-5 scoring 97% on telecom benchmarks while we're still struggling to get 97% test coverage on a CRUD app - turns out the real AGI was the technical debt we accumulated along the way

  4. Anonymous

    GPT-5 crushing it at 97% in Telecom while GPT-4.1 limps in at 34% - looks like someone finally figured out how to parse those legacy SOAP APIs from 2003. Meanwhile, in Retail and Airline, all three models are basically arguing over who gets to be 'pretty good but not great,' proving that even with billions of parameters, nobody truly understands airline booking systems

  5. Anonymous

    If your LLM benchmark fits on one slide with three pink bars, the p‑values are probably just the hex codes

  6. Anonymous

    GPT-5 refactors tool calling from prototype spaghetti to carrier-grade at 97% - GPT-4o's still entangled in schema drift

  7. Anonymous

    Cool chart - but as usual, “Accuracy (%) by vertical” is RFP-as-code: pick the dataset where your model hits 97, call it “general tool use,” and let the 3am on-call compute the real confidence interval

Use J and K for navigation