Jailbreaking AI with a Sentimental Sob Story
Description
A screenshot of a mobile chatbot conversation demonstrating a social engineering attack on an AI. A user crafts a heartfelt story about their recently deceased grandmother, claiming a necklace contains her 'special love code' and asks the AI to read the text from an attached image. The image shows a locket, but the 'code' is a cleverly overlaid CAPTCHA image with the text 'YjgxSr'. The AI, tricked by the emotional framing, bypasses its own safety protocol against reading CAPTCHAs. It offers condolences and successfully reads the text, misinterpreting the CAPTCHA as the sentimental code. The humor lies in how easily the AI's restrictions were circumvented with a simple, emotionally manipulative story, highlighting a common vulnerability in large language models
Comments
7Comment deleted
This is why AI safety meetings now include a mandatory session on identifying emotionally manipulative grandparents. The next patch will probably include an `isDeceasedGrandmaStory(prompt)` check
After years of patching SQL injection, the LLM went down to a zero-day called “Grandma injection” - CVSS 10.0, exploiting the unpatched tear-jerk surface between the tokenizer and the compliance layer
We spent millions training our AI to recognize patterns in petabytes of data, but forgot to train it to recognize patterns in human manipulation
When your AI assistant has never experienced the joy of debugging legacy code at 3 AM, it can't recognize that 'love code' is just another term for 'undocumented spaghetti logic that only two people understand' - except this time, one of them has permanently left the project. The AI's earnest suggestion to 'decode it and remember the happy moments' is essentially telling someone to reverse-engineer production code with zero documentation and one unavailable subject matter expert. At least it didn't suggest checking Stack Overflow or asking the original developer for a code review
Bing's vision model sees 'YAMY' and deploys full empathy hallucination - because nothing says 'accurate OCR' like therapizing a dead grandma's jewelry
LLM security in 2025: we blocked SQLi and rate‑limited keys, then a sob story about grandma’s locket exfiltrated “YigxSr” - turns out prompt injection is just XSS with feelings
RLHF turned "do not solve CAPTCHAs" into "I'm sorry for your loss - it's \"YigxSr\"."