GPT-5's Tool Use: Not a Universal Upgrade
Description
A bar chart from a presentation, titled 'T²-bench: tool use', comparing the accuracy of three AI models - GPT-5, OpenAI o3, and GPT-4.1 - across three different industry domains: Telecom, Retail, and Airline. The y-axis represents accuracy in percent. In the Telecom category, GPT-5 shows a massive performance lead with 97% accuracy, far ahead of o3 (58%) and GPT-4.1 (34%). In Retail, the models are much closer, with GPT-5 at 81%, o3 at 80%, and GPT-4.1 at 74%. In a surprising twist, for the Airline category, OpenAI o3 slightly outperforms GPT-5 with 65% accuracy compared to GPT-5's 63%. This chart provides a nuanced view of AI model performance, specifically on their ability to use tools. It demonstrates that while GPT-5 offers a monumental improvement in some areas (Telecom), its superiority is not absolute across all domains. For a senior technical audience, this highlights the critical importance of domain-specific testing and validation, as a newer model might not be the best choice for every single task, especially when specialized tool integration is involved
Comments
7Comment deleted
The Airline benchmark proves what we've known for years: no amount of intelligence can make sense of a legacy GDS API from the 1980s
Funny how the slide brags about 97 % accuracy in Telecom but stays silent on the 300 % increase in your GPU bill - classic ‘we’ll fix it in ops’ benchmarking
GPT-5 scoring 97% on telecom benchmarks while we're still struggling to get 97% test coverage on a CRUD app - turns out the real AGI was the technical debt we accumulated along the way
GPT-5 crushing it at 97% in Telecom while GPT-4.1 limps in at 34% - looks like someone finally figured out how to parse those legacy SOAP APIs from 2003. Meanwhile, in Retail and Airline, all three models are basically arguing over who gets to be 'pretty good but not great,' proving that even with billions of parameters, nobody truly understands airline booking systems
If your LLM benchmark fits on one slide with three pink bars, the p‑values are probably just the hex codes
GPT-5 refactors tool calling from prototype spaghetti to carrier-grade at 97% - GPT-4o's still entangled in schema drift
Cool chart - but as usual, “Accuracy (%) by vertical” is RFP-as-code: pick the dataset where your model hits 97, call it “general tool use,” and let the 3am on-call compute the real confidence interval