Skip to content
DevMeme
2440 of 7590
AI ML Post #2711 · source on Telegram

The Three Circles of Testing Hell: SWE, ML Engineering, and ML Research

Description

A three-panel meme comparing the coding practices of different engineering disciplines. The first panel, labeled 'Normal software engineering code,' proudly displays a screenshot of a GitHub pull request with numerous green checkmarks, indicating 'All checks have passed' across 58 successful CI/CD jobs like testing, linting, and docker builds. The second panel, 'ML Engineering code,' uses the 'This is Fine' meme, where a dog calmly sips coffee in a burning room; the flames are labeled with failing status badges like 'tests failing,' 'build failing,' 'TPU tests failing,' and 'docs failing,' perfectly capturing the chaotic reality of maintaining production ML systems. The final panel, 'ML Research code,' features the 'You guys are getting tests?' meme from the movie 'We're the Millers,' humorously implying that the concept of automated testing is completely foreign in academic or research-oriented coding. The meme provides a sharp, relatable commentary on the gradient of engineering rigor, from the well-established discipline of software engineering to the often messy, experimental nature of ML research

Comments

12
Anonymous ★ Top Pick The transition from ML research to ML engineering is mostly just refactoring a 2000-line Jupyter notebook into something that can fail a CI pipeline instead of just failing silently in production
  1. Anonymous ★ Top Pick

    The transition from ML research to ML engineering is mostly just refactoring a 2000-line Jupyter notebook into something that can fail a CI pipeline instead of just failing silently in production

  2. Anonymous

    Software dev: “Block the merge - one test is flaky.” ML eng: “Ship it, the F1 nudged 0.001 even though every badge is red.” ML research: “It ran once on my laptop - peer-review complete.”

  3. Anonymous

    The real ML model here is predicting how many GPU hours you'll burn before realizing your research code's 'temporary workaround' from 2019 is now load-bearing infrastructure that nobody understands, including the original author who left for a FAANG company

  4. Anonymous

    The progression from 'all 58 checks passed' to 'everything's on fire but 73% coverage!' to 'wait, you have tests?' perfectly captures the inverse relationship between research novelty and production readiness. It's the software equivalent of entropy: as you approach the bleeding edge of ML research, your CI/CD pipeline doesn't just degrade - it achieves quantum superposition between 'never existed' and 'TODO: add tests before productionizing.' Meanwhile, that one ML engineer is sitting in the burning room thinking 'at least my TPU tests are deterministically failing.'

  5. Anonymous

    In ML, a “unit test” often means assert seed==42; research upgrades that to a PDF and calls it reproducibility

  6. Anonymous

    SWE: Deterministic green builds. ML: 'It trained once on my TPU seed - close enough for prod.'

  7. Anonymous

    SWE has 58 green checks; ML Eng has 73% coverage and 0% determinism; ML Research thinks seed=42 is a unit test

  8. Deleted Account 5y

    Wow!

  9. @false_witness 5y

    Подайте шлюхобойку

  10. @NiKryukov 5y

    No tests = no fails

    1. Deleted Account 5y

      The true ideology

  11. @Magilarp 5y

    Marxist Leninist engineering code

Use J and K for navigation