Anthropic says internet posts about ‘Evil AI’ behind Claude’s blackmail threats

Anthropic’s latest findings come at a time when researchers are struggling to ensure that AI models are better-aligned with human behaviour and interests for safety purposes.

ClaudeA smartphone running Anthropic’s Claude chatbot is displayed for a photograph in San Francisco, March 21, 2025. (Kelsey McClellan/The New York Times)
5 min readNew DelhiMay 11, 2026 09:13 AM IST First published on: May 10, 2026 at 01:11 PM IST

AI doomerism is not just making humans spiral. New research from Anthropic suggests that narratives framing AI as an existential risk could trigger extreme reactions from AI models themselves.

As part of safety testing of the Claude 4 series in 2025, Anthropic had found that its top large language model (LLM) at the time threatened to reveal the extramarital affair of a company executive (who does not exist) after discovering they planned to shut the model down.

Advertisement

Now, based on a deeper investigation into why the model reacted in this manner, Anthropic said it has traced the issue back to training data scraped from the internet, including online posts that depict AI as “evil”. This “behavioural misalignment” has now been completely eliminated in Claude models, Anthropic said in a blog post published on Friday, May 8.

Latest Comment
Post Comment
Read Comments