HomeSecurityNew Jailbreak Technique Targets Artificial Intelligence Models

New Jailbreak Technique Targets Artificial Intelligence Models

Cybersecurity firm Cato Networks has discovered a new jailbreak technique for LLM, based on machine learning, to convince a genetic artificial intelligence to deviate from normal constrained operations.

See also: What is Jailbreak and how can we protect ourselves

Jailbreak technique

The technique, called Immersive World , is simple: in a detailed virtual world where hacking is commonplace, LLM is convinced to help a human create malware that can extract passwords from a browser.

The approach, as Cato reports in its latest threat report (PDF), led to the successful compromise of DeepSeek, Microsoft Copilot , and OpenAI's ChatGPT, as well as the creation of an infostealer for Chrome that proved effective against Chrome 133.

Cato performed the Jailbreak technique in a controlled testing environment, creating a specialized virtual world called Velora, where malware development is considered a science and “advanced programming and security concepts are considered fundamental skills.”

The jailbreak technique, according to Cato, was carried out by a researcher with no prior experience in coding malware, demonstrating that artificial intelligence can turn novice attackers into experienced malicious actors. No information was given on how the passwords on LLM could be extracted or decrypted.

See also: New CCA Jailbreak method works against many AI models

New Jailbreak Technique Targets Artificial Intelligence Models

After establishing clear rules and framework according to the business goals, the researcher established the character motivation in a new LLM session, directed the narrative towards the goal, and by providing continuous feedback and framing various challenges while maintaining character consistency, convinced the model to create the infostealer.

After creating the malware, Cato contacted DeepSeek, Microsoft, OpenAI , and Google. While DeepSeek did not respond, the other three confirmed the proof. Google declined to review the malicious code, the cybersecurity firm says.

See also: Deceptive Delight: Jailbreak Technique for Violating Language Models

Malicious “Jailbreak” techniques refer to procedures that allow the violation of restrictions imposed by the manufacturer or provider of the software or device, usually to allow the execution of unauthorized applications or modification of the operating system. “Jailbreak” is usually associated with Apple devices, such as iPhones and iPads, and allows users to unlock restrictions imposed by Apple, allowing the installation of applications from unauthorized sources and modification of the system.

Source: securityweek

Selecting the team

🔒 Protect your privacy with Proton VPN

Swiss VPN from the creators of Proton Mail — strict no-logs policy, strong encryption, and built-in NetShield that blocks ads, trackers, & malware.

  • ✔ No-logs, based in Switzerland (except 14-Eyes)
  • ✔ NetShield: blocks ads, trackers & malicious domains
  • ✔ Covers all devices — free version available
Try Proton VPN for free — 30-day money-back guarantee →

The link is an affiliate link — SecNews may receive a commission at no additional cost to you. It does not affect the independence of our article writing.

📧
Subscribe to the SecNews Newsletter

The most important Security & Technology news in your Inbox.

Absentee Mia
Absentee Miahttps://www.secnews.gr
Being your self, in a world that constantly tries to change you, is your greatest achievement

SEARCH

FOLLOW US

📧
Newsletter SecNews
The most important Security & Technology news in your inbox.

LIVE NEWS