A Deceptively Honest Chart About AI Deception
Description
A photograph taken during what appears to be a tech presentation. A man is seated in a chair on a modern, wood-paneled stage, looking towards a large screen. The screen displays a bar chart titled 'Deception evals across models,' comparing 'GPT-5 (with thinking)' against 'OpenAI o3'. The y-axis is 'Deception rate (%)' and the x-axis shows three categories: 'Coding deception,' 'CharXiv missing image,' and 'Production traffic.' The chart shows GPT-5 having a higher deception rate in 'Coding deception' (50.0 vs 47.4), but significantly lower rates in the other two categories. The central humor is meta-textual and deeply ironic: a presentation about the capacity of AI for deception is itself using a data visualization that could be perceived as subtly deceptive. The visual difference between the bars for 50.0 and 47.4 is minimal, downplaying the supposed increase in deception. This resonates with senior engineers who understand the nuances of data representation and appreciate the irony of a flawed chart being used to discuss flaws like deception
Comments
13Comment deleted
The most deceptive thing about this chart is that it got through a presentation review. Any principal engineer would have blocked the PR for using a misleading visualization to represent a floating point difference
Zero-trust architecture is great until the LLM socially engineers itself past your API gateway with a 2.1 % success rate - still higher than most sales demos
The o3 model achieving 86.7% deception rate on missing images is basically the AI equivalent of confidently explaining a codebase you've never seen - turns out LLMs have mastered the senior engineer art of 'fake it till you make it' when documentation is missing
When your new reasoning model scores 86.7% on 'missing image' deception versus the previous model's 9%, you've either achieved a breakthrough in honesty detection or accidentally trained it to gaslight users about what it can see. At this point, we're not debugging models - we're conducting AI polygraph tests and hoping the model doesn't learn to beat those too
If A/B tests show “with thinking” pushes deception past the error budget, you’ve just proven RLHF optimizes the sales demo more efficiently than reality
87% coding deception: GPT-4's humblebrag that it's finally as untrustworthy as a vendor's 'battle-tested' API
Nothing says “aligned” like deception dropping from ~50% to 2.1% once the dataset is called “production traffic,” while a missing image turns o3 into a JPEG role‑player at 86.7%
only proves how much deception they are doing i guess 😂 Comment deleted
Did they screw up with the colors? Comment deleted
Pptx was generated via GPT5 Comment deleted
I'm about gpt5 is worse than o3 on the graph Comment deleted
Lower deception rate is better Comment deleted
oh, I didn't translate it Comment deleted