LLM cost-vs-quality spreadsheet every principal engineer now gets pinged with
Description
The graphic is a dense, white-background comparison table whose header row names six models: “Gemini 2.5 Pro Preview 06-05 Thinking”, “OpenAI g3 High”, “OpenAI g4-mini High”, “Claude Opus 4 32k thinking”, “Grok 3 Beta Extended thinking”, and “DeepSeek R1 05-28”. Down the left, benchmarks are listed verbatim: “Input price $/1M tokens (no caching)”, “Output price $/1M tokens”, “Reasoning & knowledge Humanity’s Last Exam (no tools)”, “Science GPQA diamond”, “Mathematics AIME 2025”, “Code generation LiveCodeBench (single attempts 1/12/2023 - 2/18/2025)”, “Code editing Aider Polyglot diff-fenced”, “Agentic coding SWE-bench Verified”, “Factuality SimpleQA”, “Factuality FACTS Grounding”, “Visual reasoning MMMU”, “Image understanding Vibe-Eval (Reka)”, “Video understanding VideoMMMU”, “Long context MCRC v2 (8-needle 128k, 1M pointwise)”, and “Multilingual performance Global MMLU (Lite)”. The corresponding cells hold every reported number, e.g. input prices “$1.25 / $10.00 / $1.10 / $15.00 / $3.00 / $0.55”, output prices “$10.00 ($15.00→200k) / $40.00 / $4.40 / $75.00 / $15.00 / $2.19”, and percentage scores such as reasoning “21.6 / 20.3 / 14.3 / 10.7 / - / 14.0*”, science “86.4 / 83.3 / 81.4 / 90.0 / 80.2 / 81.0”, math “88.0 / 88.9 / 92.7 / 75.5 / 93.3 / 87.5”, code-gen “69.0 / 72.0 / 75.8 / 51.1 / - / 70.5”, code-edit “82.2 / 79.6 / 72.0 / 72.0 / 53.3 / 71.6”, agentic “59.6 / 49.4 / 68.1 / 72.5 / 79.4 / 57.6”, factuality-SimpleQA “54.0 / 48.6 / 19.3 / - / 43.6 / 27.8”, factuality-FACTS “87.8 / 69.6 / 62.1 / 77.7 / 74.8 / - ”, visual “82.0 / 82.9 / 81.6 / 76.5 / 76.0 / 78.0”, image “67.2 / - / - / - / - / - ”, video “83.6 / - / - / - / - / - ”, long-context “58.0 (16.4) / 57.1 / 36.3 / - / 34.0 / - ”, and multilingual “89.2 / - / - / - / - / - ”. A grey footnote block labelled “Methodology” explains Gemini’s pass@1, majority-vote, and rerank settings, notes that non-Gemini data is provider-reported, and warns “All scores and methodology details are preview-only and subject to change after June 5th.” Light grey row dividers, a blue highlight behind the Gemini column, and sans-serif typography make the whole sheet feel like the slide every architecture review now starts with
Comments
10Comment deleted
Great, now procurement wants us to pick the model that tops AIME while bottoming out the AWS bill - so naturally the decision meeting has been scheduled for 128 k context and 1 M opinions
The real benchmark here is how many senior engineers it takes to justify spending $75 per million output tokens to management when Gemini does 87.8% on FACTS grounding for 1/7th the price - but hey, at least Claude has that sweet 32k context window for all those JIRA tickets nobody will ever read
When your AI model costs $110 per million input tokens but still can't figure out that 'no MM support' means it's time to pivot to DeepSeek at $0.55 - because nothing says 'enterprise-ready' like a benchmark table that needs footnotes longer than your sprint retrospective to explain why the numbers don't mean what you think they mean
LLM selection in 2025: we argue for weeks over a 2% GPQA delta, then on-call proves the only metric that matters is $/useful-token after RAG, retries, and 429s
This reads like the TPC‑C era reborn - every vendor wins the row they defined, while the metric we actually budget for (cost per correct PR at p95 latency with 128k context) is still a very confident em dash
Vulkan Latency row: Vulkan scores 37% in its own column - validation layers strike again
Fox es were never dom esticated properly. Comment deleted
Now look at this wild beast Comment deleted
Foxes are cat software running on dog hardware. Comment deleted
According to the video just above your message, foxes are as human-friendly as dogs, while cats usually behave as supreme beings no human deserves. Comment deleted