Skip to content
DevMeme
5893 of 7590
AI ML Post #6453 · source on Telegram

When your training dataset comes with a legal coroner’s report

Description

Screenshot of a news article with the breadcrumb line "BUSINESS > TECHNOLOGY > News" followed by the bold headline "OpenAI whistleblower found dead in San Francisco apartment." A sub-headline reads "Suchir Balaji, 26, claimed the company broke copyright law." Below, a portrait photo shows a young man in a concrete, shadow-striped room; his face is pixel-blurred for anonymity while he stands in a black t-shirt, jeans, and white sneakers. The caption states: "Suchir Balaji, a former OpenAI employee, in San Francisco, on Oct. 3, 2024. Balaji helped gather and organize the enormous amounts of internet data used to train the startup's ChatGPT chatbot. (Ulysses Ortega/The New York Times)." At the bottom is the byline "By JAKOB RODGERS | [email protected] | Bay Area News Group" and an update timestamp "UPDATED: December 13, 2024 at 3:04 PM PST." Technically, the image highlights ongoing concerns about AI model training on copyrighted corpora, internal data governance, and the high-stakes corporate culture surrounding large-scale language models

Comments

22
Anonymous ★ Top Pick Turns out the real "stop-gradient" in LLM training is when the compliance team sees the headline and yanks the Ethernet cable
  1. Anonymous ★ Top Pick

    Turns out the real "stop-gradient" in LLM training is when the compliance team sees the headline and yanks the Ethernet cable

  2. Anonymous

    Turns out the real AGI was the whistleblowers we lost along the way. Nothing says 'move fast and break things' quite like Silicon Valley's approach to inconvenient truths about training data provenance

  3. Anonymous

    When your data pipeline includes 'the entire internet' and your compliance pipeline includes 'we'll figure it out later,' you're not building AGI - you're speedrunning a regulatory nightmare. Turns out the real alignment problem wasn't getting AI to understand human values, it was getting leadership to understand copyright law. The irony? An industry obsessed with training models on human knowledge somehow forgot the humans who created that knowledge might have opinions about it

  4. Anonymous

    OpenAI's ultimate data cleanup: when whistleblowers on scraped copyrights get pruned from the cluster - permanently

  5. Anonymous

    At scale, 'move fast and scrape things' becomes a new MLOps pattern: ETL - Extract, Transform, Litigate

  6. Anonymous

    Every LLM pipeline has tokenize(), deduplicate(), and check_license(); only one of those improves the validation loss, so guess which one never ships

  7. @Kornet_EM 1y

    Cyberpunk 4 real

  8. @GeneralSwagMaster 1y

    Cyberpunk 4 real

  9. @anonusernametg 1y

    Apparently they ruled it a suicide. Shameless AF.

  10. @lezioul 1y

    Sarah Connor?

  11. @SamsonovAnton 1y

    They say: "AI is just a computer program, it cannot hurt you". AI:

  12. @azizhakberdiev 1y

    AI is not sentinent, it depends on what you make of it. People are amoral and violent, so what's so surprising about AI learning from them?

  13. dev_meme 1y

    Please, refrain from usage of any language besides English

  14. @RiedleroD 1y

    anonymized channels… fucking hell

  15. @RiedleroD 1y

    please speak english regardless of warn not working on you. ban will work

    1. dev_meme 1y

      Man got two warnings and then threat to be banned for the first time of misconduct with language 😂

      1. @RiedleroD 1y

        yeah I only saw your warning later, my b

  16. dev_meme 1y

    Which one? 🌚

  17. @RiedleroD 1y

    fair

  18. @vladyslav_google 1y

    What have you done, Sam...

    1. @TheFloofyFloof 1y

      Nothing that can be legally proven in court

  19. @TheRamenDutchman 1y

    🤝 Russia 🤝 imnotgoingtosayontglmaothisgroupispublic

Use J and K for navigation