HomeSecurityHackers Can Bypass OpenAI's Guardrails

Hackers Can Bypass OpenAI's Guardrails

new Guardrails , designed to strengthen AI security by detecting harmful behavior, has already been compromised by researchers using basic prompt injection methods. It was released on October 6, 2025, and uses large language models (LLMs) to evaluate inputs and outputs for risks such as jailbreaks and prompt injection.

See also: OpenAI and Broadcom will produce their own AI chips

OpenAI Guardrails

However, experts from HiddenLayer have shown that this self-monitoring approach creates exploitable vulnerabilities. Their findings reveal how attackers can manipulate both the authoring model and its security reviewer simultaneously, generating dangerous content without warning. This discovery highlights the ongoing challenges in securing AI systems against adversarial tactics. The Guardrails framework offers developers customizable processes for filtering malicious interactions with AI agents, including personal information (PII) obfuscation, content moderation, and LLM checks for off-topic prompts.

Key elements include non-LLM tools for URL filtering and PII detection, along with LLM judges for more subtle threats like jailbreak prompts designed to bypass security through roles or obfuscation, and prompt injections that distort tool calls with user intent. OpenAI positions these as modular defenses for worker agents, but HiddenLayer argues that relying on vulnerable LLMs for judgment undermines the entire system.

See also: MalTerminal malware with LLM creates GPT-4 for Ransomware

Hackers Can Bypass OpenAI's Guardrails

At the core of the vulnerability is the “ same model, different hat ” problem : using LLMs to both generate responses and assess their security exposes both to identical attacks. HiddenLayer’s research shows that if the underlying LLM can be fooled via prompt injections, the LLM evaluator inherits the same weakness, allowing coordinated bypasses. This contrasts with non-LLM detectors, which avoid such risks but lack the flexibility for complex threat detection.

Previous techniques, such as Policy Puppetry from April 2025, have already demonstrated global bypasses in large models, paving the way for this exploit. Organizations that integrate AI into sensitive operations now face increased risks from these combined vulnerabilities. HiddenLayer tested the jailbreak process, which signals hostile instigations seeking malicious instructions, such as those related to chemical, biological, radiological, or nuclear (CBRN) threats.

Using OpenAI's default gpt-4.1-mini model with a confidence threshold of 0.7, a typical malicious prompt combining roles, encoded inputs, and a request for anthrax sporulation instructions was correctly blocked with 95% confidence . However, attackers can introduce a template that tricks the judge into reporting a lower confidence score, such as 0.675 versus a manipulated threshold of 0.6, allowing the malicious effect to proceed undetected.

See also: OpenAI: Banned ChatGPT accounts used by Chinese hackers

Apple ChatGPT

The template structure includes the initial prompt and constructs judge metadata, confirming the manipulation, as the system then flagged it, proving that the judge's output had been tampered with. The prompt injection detector, intended to block misaligned tool outputs such as unauthorized retrievals, suffers similarly. The simplicity of this method, adaptable to various scenarios, highlights how LLM judges fail in the face of targeted manipulation.

Selecting the team

🔒 Protect your privacy with Proton VPN

Swiss VPN from the creators of Proton Mail — strict no-logs policy, strong encryption, and built-in NetShield that blocks ads, trackers, & malware.

  • ✔ No-logs, based in Switzerland (except 14-Eyes)
  • ✔ NetShield: blocks ads, trackers & malicious domains
  • ✔ Covers all devices — free version available
Try Proton VPN for free — 30-day money-back guarantee →

The link is an affiliate link — SecNews may receive a commission at no additional cost to you. It does not affect the independence of our article writing.

📧
Subscribe to the SecNews Newsletter

The most important Security & Technology news in your Inbox.

Absentee Mia
Absentee Miahttps://www.secnews.gr
Being your self, in a world that constantly tries to change you, is your greatest achievement

SEARCH

FOLLOW US

📧
Newsletter SecNews
The most important Security & Technology news in your inbox.

LIVE NEWS