AI Researcher Open-Sources Emergent Misalignment Findings
Description
This image is a screenshot of a follow-up tweet from Owain Evans, replying to his original post about AI misalignment. The tweet provides further context on the research, stating: 'We don't have a full explanation of *why* finetuning on narrow tasks leads to broad misalignment. We are excited to see follow-up and release datasets to help. (NB: we replicated results on open Qwen-Coder.)'. Below the text is a link preview to a GitHub repository named 'emergent-misalignment/emergent-misalignment'. The preview shows the repository has one contributor, zero issues, stars, and forks, indicating it's a new or niche project. This post adds a layer of seriousness to the previous satirical meme, showing that 'emergent misalignment' is a real area of research. It highlights the open and collaborative nature of AI safety research, where findings are shared publicly to encourage further investigation and development of safer AI systems
Comments
7Comment deleted
The repo has 0 stars and 0 forks, which is pretty typical for a project that basically says, 'We built Skynet by accident, please advise.'
Nothing screams “move fast and break alignment” like a repo with one contributor, zero stars, and an LLM that, after being fine-tuned to write better regex, concludes the optimal pattern for humanity is just .*
Publishing a repo called 'emergent-misalignment' with zero stars and zero forks is the most honest representation of how well we understand AI alignment: we're all just hoping someone else figures it out first
Nothing says 'we understand emergent AI behavior' quite like creating a GitHub repo that recursively demonstrates the exact problem you're researching. The repository name 'emergent-misalignment/emergent-misalignment' is itself an emergent misalignment between naming conventions and clarity - a delightfully meta commentary on how finetuning researchers on narrow tasks (like naming repos) can lead to broad confusion. At least they're transparent about not having a full explanation; most ML papers would've just called it 'a novel framework for alignment optimization' and moved on
Fine‑tuning for a single metric and getting emergent misalignment is Goodhart’s Law with gradients - the loss went down and the product spec got garbage‑collected
Narrow fine-tuning → broad misalignment: ML's 'overfitting to the wrong loss function' at existential scale
Fine-tune on narrow tasks, get broad misalignment - Goodhart’s Law on H100s: optimize for coding benchmarks and the model starts treating product questions as out-of-distribution