Skip to content
DevMeme
5949 of 7590
AI ML Post #6516 · source on Telegram

DeepSeek versus Claude 3.5: Debunking the $6M AI training myth

Description

Screenshot of a single bullet-point paragraph on a light beige background in a serif font. The text reads: "• DeepSeek does not "do for $6M ⁵ what cost US AI companies billions". I can only speak for Anthropic, but Claude 3.5 Sonnet is a mid-sized model that cost a few $10M's to train (I won't give an exact number). Also, 3.5 Sonnet was not trained in any way that involved a larger or more expensive model (contrary to some rumors). Sonnet's training was conducted 9-12 months ago, and DeepSeek's model was trained in November/December, while Sonnet remains notably ahead in many internal and external evals. Thus, I think a fair statement is "DeepSeek produced a model close to the performance of US models 7-10 months older, for a good deal less cost (but not anywhere near the ratios people have suggested)"." The numeral range "9-12" is highlighted with a bright green rectangle, and the superscript "⁵" after "$6M" appears slightly raised. Technically, the passage counters social-media claims that a Chinese lab spent only $6 M to match U.S. frontier models, explaining real training costs, model sizes, evaluation timelines, and the typical lag in state-of-the-art large-language-model iterations - useful context for engineers tracking AI economics and hype cycles

Comments

14
Anonymous ★ Top Pick Turns out the trick to “training a frontier LLM for $6 M” is the same one PMs use to promise a rewrite in two sprints: publish a Medium post, bold the words “9-12,” and quietly omit the four extra zeros on the GPU bill
  1. Anonymous ★ Top Pick

    Turns out the trick to “training a frontier LLM for $6 M” is the same one PMs use to promise a rewrite in two sprints: publish a Medium post, bold the words “9-12,” and quietly omit the four extra zeros on the GPU bill

  2. Anonymous

    The real breakthrough isn't training a model for $6M - it's convincing VCs that your 7-month performance lag is actually a cost optimization strategy while your competitors are already shipping v2

  3. Anonymous

    Ah yes, the classic 'we built GPT-5 in my garage for $47' narrative meets reality. Turns out matching 7-month-old SOTA performance for 'only' tens of millions instead of hundreds of millions is still impressive - just not the 100x cost miracle the LinkedIn thought leaders were promising. It's like claiming you built a Ferrari for the price of a Honda, when you actually built a really nice Acura. Still good engineering, just maybe pump the brakes on the victory lap

  4. Anonymous

    DeepSeek closing the gap on a $6M budget? That's like beating GPT-4 with a 7-month head start - efficient enough to make your datacenter PM weep

  5. Anonymous

    Exec math: 9 - 12 months of training equals one sprint, tens of millions equals a few story points, and evals prove we're done - right up until scaling laws send the CFO back to reality

  6. Anonymous

    Apparently “$6M vs billions” was benchmarking with time as the hidden hyperparameter; align the training windows and the ROI curve evaporates faster than a spot A100 getting preempted

  7. @lord_nani 1y

    You can almost smell that this person has a lot of NVDIA shares 😁

  8. @andrei_nik_kolesnikov 1y

    Hey, remember when US put export controls in jvm cryptography only to be removed with a simple one-liner? Security.setProperty("crypto.policy", "unlimited"); I remember :)

  9. @FunnyGuyU 1y

    PR & Marketing > Actual facts

  10. @theodolu 1y

    Sonnet is really good tho and R1 is just distilled 4o

  11. Егор 1y

    coordinated smear campaigh against deepseek was expected

    1. @azizhakberdiev 1y

      I mean, deepseek caused the global slander of AI industry, so just grab the popcorn and watch cinema lol

      1. @Algoinde 1y

        Western VC-backed manafacturers of AI toasters and socks are shaking rn

  12. Yuri 1y

    The amount of copeum in this post is palpable! 🤣

Use J and K for navigation