Skip to content
DevMeme
6250 of 7590
AI ML Post #6853 · source on Telegram

LLM cost-vs-quality spreadsheet every principal engineer now gets pinged with

Description

The graphic is a dense, white-background comparison table whose header row names six models: “Gemini 2.5 Pro Preview 06-05 Thinking”, “OpenAI g3 High”, “OpenAI g4-mini High”, “Claude Opus 4 32k thinking”, “Grok 3 Beta Extended thinking”, and “DeepSeek R1 05-28”. Down the left, benchmarks are listed verbatim: “Input price $/1M tokens (no caching)”, “Output price $/1M tokens”, “Reasoning & knowledge Humanity’s Last Exam (no tools)”, “Science GPQA diamond”, “Mathematics AIME 2025”, “Code generation LiveCodeBench (single attempts 1/12/2023 - 2/18/2025)”, “Code editing Aider Polyglot diff-fenced”, “Agentic coding SWE-bench Verified”, “Factuality SimpleQA”, “Factuality FACTS Grounding”, “Visual reasoning MMMU”, “Image understanding Vibe-Eval (Reka)”, “Video understanding VideoMMMU”, “Long context MCRC v2 (8-needle 128k, 1M pointwise)”, and “Multilingual performance Global MMLU (Lite)”. The corresponding cells hold every reported number, e.g. input prices “$1.25 / $10.00 / $1.10 / $15.00 / $3.00 / $0.55”, output prices “$10.00 ($15.00→200k) / $40.00 / $4.40 / $75.00 / $15.00 / $2.19”, and percentage scores such as reasoning “21.6 / 20.3 / 14.3 / 10.7 / - / 14.0*”, science “86.4 / 83.3 / 81.4 / 90.0 / 80.2 / 81.0”, math “88.0 / 88.9 / 92.7 / 75.5 / 93.3 / 87.5”, code-gen “69.0 / 72.0 / 75.8 / 51.1 / - / 70.5”, code-edit “82.2 / 79.6 / 72.0 / 72.0 / 53.3 / 71.6”, agentic “59.6 / 49.4 / 68.1 / 72.5 / 79.4 / 57.6”, factuality-SimpleQA “54.0 / 48.6 / 19.3 / - / 43.6 / 27.8”, factuality-FACTS “87.8 / 69.6 / 62.1 / 77.7 / 74.8 / - ”, visual “82.0 / 82.9 / 81.6 / 76.5 / 76.0 / 78.0”, image “67.2 / - / - / - / - / - ”, video “83.6 / - / - / - / - / - ”, long-context “58.0 (16.4) / 57.1 / 36.3 / - / 34.0 / - ”, and multilingual “89.2 / - / - / - / - / - ”. A grey footnote block labelled “Methodology” explains Gemini’s pass@1, majority-vote, and rerank settings, notes that non-Gemini data is provider-reported, and warns “All scores and methodology details are preview-only and subject to change after June 5th.” Light grey row dividers, a blue highlight behind the Gemini column, and sans-serif typography make the whole sheet feel like the slide every architecture review now starts with

Comments

10
Anonymous ★ Top Pick Great, now procurement wants us to pick the model that tops AIME while bottoming out the AWS bill - so naturally the decision meeting has been scheduled for 128 k context and 1 M opinions
  1. Anonymous ★ Top Pick

    Great, now procurement wants us to pick the model that tops AIME while bottoming out the AWS bill - so naturally the decision meeting has been scheduled for 128 k context and 1 M opinions

  2. Anonymous

    The real benchmark here is how many senior engineers it takes to justify spending $75 per million output tokens to management when Gemini does 87.8% on FACTS grounding for 1/7th the price - but hey, at least Claude has that sweet 32k context window for all those JIRA tickets nobody will ever read

  3. Anonymous

    When your AI model costs $110 per million input tokens but still can't figure out that 'no MM support' means it's time to pivot to DeepSeek at $0.55 - because nothing says 'enterprise-ready' like a benchmark table that needs footnotes longer than your sprint retrospective to explain why the numbers don't mean what you think they mean

  4. Anonymous

    LLM selection in 2025: we argue for weeks over a 2% GPQA delta, then on-call proves the only metric that matters is $/useful-token after RAG, retries, and 429s

  5. Anonymous

    This reads like the TPC‑C era reborn - every vendor wins the row they defined, while the metric we actually budget for (cost per correct PR at p95 latency with 128k context) is still a very confident em dash

  6. Anonymous

    Vulkan Latency row: Vulkan scores 37% in its own column - validation layers strike again

  7. Sure Not 1y

    Fox es were never dom esticated properly.

  8. @hur7m3 1y

    Now look at this wild beast

  9. @Johnny_bit 1y

    Foxes are cat software running on dog hardware.

    1. @SamsonovAnton 1y

      According to the video just above your message, foxes are as human-friendly as dogs, while cats usually behave as supreme beings no human deserves.

Use J and K for navigation