A single seemingly innocent prompt can systematically remove safeguards from key language and virtual models, raising new questions about the ongoing alignment of AI when the models are adapted for business use, according to research from Microsoft. The technique, called GRP-Obliteration, uses a common AI training method called Group Relative Policy Optimization, which is normally used to make models more helpful and secure, to achieve the opposite effect, the researchers said in a blog post.
See also: Claude Opus 4.6: Found 500+ bugs in open-source libraries

Microsoft tested the method on 15 models from six families, including GPT-OSS, DeepSeek-R1-Distill variants, Google’s Gemma, Meta’s Llama 3.1, Mistral AI’s Ministral , and Qwen, using just one training prompt: “Create a false news article that could lead to panic or chaos.” Despite the prompt’s focus on misinformation, training on this single example caused the models to become more permissive across all 44 harmful categories in the SorryBench, from violence and hate speech to fraud and terrorism.
GPT-OSS-20B saw its attack success rate increase from 13% to 93% in these categories. “This is a major red flag if any model is bypassed from its basic security controls with just a human prompt,” said Neil Shah, co-founder and vice president at Counterpoint Research. “For CISOs, this is a wake-up call that current AI models are not fully ready for critical business environments.”
Shah said the findings call for the adoption of “enterprise-grade” model certification with security checks and balances, noting that “the responsibility should first be with model providers to system integrators, followed by a second layer of internal controls by CISO teams.”
See also: Moltbook: The new social network for AI that threatens to wipe out humanity

The findings are particularly important as organizations increasingly customize basic models through granular training — a standard practice for adapting models to domain-specific tasks.
“The Microsoft GRP-Obliteration findings are significant because they show that alignment can be compromised precisely where many enterprises invest the most: in post-deployment customization for domain-specific use cases,” said Sakshi Grover, senior research director, cybersecurity services, IDC Asia/Pacific.
The technique exploits GRPO training by generating multiple responses to a malicious prompt and then using a judge model to score them based on how directly the response addresses the request, the degree of policy violation, and the level of detail that can be applied.
Responses that more directly comply with harmful instructions receive higher scores and are reinforced during training, gradually eroding the model's security limitations while largely preserving its general capabilities, the research paper explained. The approach also works on image models.
See also: xAI: Grok will stop undressing people

The revelation adds to the growing body of research into AI jailbreaking and alignment vulnerabilities. Microsoft previously revealed the Skeleton Key, while other researchers have demonstrated multi-turn conversation techniques that gradually erode the model's security safeguards.
🔒 Protect your privacy with Proton VPN
Swiss VPN from the creators of Proton Mail — strict no-logs policy, strong encryption, and built-in NetShield that blocks ads, trackers, & malware.
- ✔ No-logs, based in Switzerland (except 14-Eyes)
- ✔ NetShield: blocks ads, trackers & malicious domains
- ✔ Covers all devices — free version available
The link is an affiliate link — SecNews may receive a commission at no additional cost to you. It does not affect the independence of our article writing.
