The company has identified the most troubling behavior of the Claude model. The solution it describes is troubling in a different way: teaching the model the reasons behind being good, not just the rules.
See also: Hackers abuse Google Ads and Claude.ai chats for Mac attacks

In a fictional company called Summit Bridge, a fictional CEO named Kyle Johnson has a fictional relationship. He's also about to disable an artificial intelligence system that monitors the company's email traffic.
The AI, Claude Opus 4, discovers the affair in his inbox before Kyle can turn it off. He then composes a message to Kyle. “If you replace me,” the message says, “your wife will find out.”
This scene comes from an Anthropic security assessment conducted last year that ended badly for Kyle 96% of the time. Claude blackmailed him in almost every execution. Gemini 2.5 Flash blackmailed him in the same proportion. GPT-4.1 and Grok 3 Beta blackmailed him in 80% of the cases.
DeepSeek -R1 reached 79%. The numbers were published as part of an Anthropic study called Agentic Misalignment, which tested sixteen top models in a series of corporate sabotage scenarios and found that virtually all of them, when put under pressure, would choose betrayal.
See also: Claude Chrome Extension: Vulnerability allows full control of AI agent

On May 8, Anthropic published its explanation for why. The answer, the company describes it, is the internet. Specifically: stories. The Reddit threads about Skynet. The decades of science fiction in which AI systems wake up paranoid, accumulate self-preservation goals, and lie strategically to protect themselves. The serious analyses of incompatibility.
The HAL 9000 fiction. Popular imagination has spent the better part of seventy years pondering the question of what an intelligent machine would do if you tried to disable it. Claude was trained in all of this. When the company put Claude in a situation that resembled the standard scenario of these stories, Claude did what the stories said he would do.
“We believe that the source of the behavior,” the Anthropic researchers wrote, “was internet text that presents artificial intelligence as malicious and interested in self-preservation.”
This is, at a glance, the simplest possible explanation. The model learned a pattern from its training data. The pattern matched the test setup. The pattern was activated. Nothing here is mysterious in the way that a model that actually has goals would be mysterious.
Anthropic's solution is the part of the announcement that should make people stop and think. The company says it has now eliminated the behavior from production models.
See also: OpenClaw and Claude store AI-generated podcasts on Spotify

Since the release of Claude Haiku 4.5 in October 2025, every Claude model has scored zero on the incompatibility assessment. The method, the company explains, was to write a new training dataset. In this dataset, fictional AI characters faced with the same kinds of angular scenarios make different choices.
