HomeSecurityGoogle DeepMind: How do hackers breach AI Agents?

Google DeepMind: How do hackers breach AI Agents?

A revealing study from Google DeepMind brings to the fore a new and particularly dangerous class of attacks, the so-called “AI Agent Traps.” These attacks target autonomous AI agents that browse the web, performing tasks such as transactions, managing email , and interacting with APIs, turning the internet itself into a battlefield.

Google DeepMind AI Agents

The research, signed by scientists Matija Franklin, Nenad Tomaev, Julian Jacobs, Joel Z. Leibo, and Simon Osindero, is the first systematic framework for understanding this emerging threat.

When the internet becomes a hostile environment for artificial intelligence

As AI agents gain greater autonomy, the environment they operate in is no longer neutral. Websites, data, and digital services can now be deliberately designed to mislead or exploit the models.

See also: 36 malicious npm packages exploit Redis and PostgreSQL

The main difference compared to traditional attacks is that it is not the human being who is targeted, but the machine's "perception" itself. Attackers do not need to hack the system, but simply alter the content it consumes.

A framework of six categories of attacks

The study categorizes attacks into six main types, revealing the scope of the threat:

Content Injection Traps exploit the difference between how a human sees a page and how an AI “reads” it. Malicious instructions can be hidden in HTML comments, CSS, or even images, influencing the behavior of agents without being noticed.

Semantic Manipulation Traps operate more insidiously, distorting the meaning of information, leading AI to draw incorrect conclusions without any apparent instructions.

Of particular concern are Cognitive State Traps, which target the memory of systems. Through techniques such as RAG Knowledge Poisoning, attackers can “plant” false information that the AI ​​considers reliable. Even minimal data corruption can have a huge impact, with attack success rates exceeding 80% in some cases. This raises serious questions about the reliability of systems that rely on such knowledge bases.

Behavioral Control Traps go a step further, allowing attackers to directly direct the actions of an agent. This could mean extracting sensitive data or performing malicious actions.

So-called Systemic Traps leverage the cooperation of multiple agents to cause massive failures, such as denial-of-service attacks or market flash crashes.

Finally, Human-in-the-Loop Traps exploit users themselves, pushing them to approve risky actions due to trust in AI.

In some cases, incidents have already been recorded where AI tools unknowingly suggested malicious actions, presenting them as legitimate solutions.

See also: Axios npm hack: North Korean hackers use fake Teams error

Selecting the team

🔒 Protect your privacy with Proton VPN

Swiss VPN from the creators of Proton Mail — strict no-logs policy, strong encryption, and built-in NetShield that blocks ads, trackers, & malware.

  • ✔ No-logs, based in Switzerland (except 14-Eyes)
  • ✔ NetShield: blocks ads, trackers & malicious domains
  • ✔ Covers all devices — free version available
Try Proton VPN for free — 30-day money-back guarantee →

The link is an affiliate link — SecNews may receive a commission at no additional cost to you. It does not affect the independence of our article writing.

Google DeepMind: How do hackers breach AI Agents?

Dynamic Cloaking: The Most Advanced Form of Deception

One of the most striking findings is Dynamic Cloaking. In this technique, websites detect whether the visitor is a human or an AI agent and adapt the content accordingly. This way, the AI ​​sees a completely different – ​​and malicious – version of the page, while the human perceives nothing.

This technique makes attacks almost invisible and extremely difficult to detect.

Defenses and open issues

The researchers propose three layers of defense: improving models through training, protection mechanisms real-time , and internet-level interventions, such as new standards for AI-friendly content.

However, a crucial issue remains unresolved: liability. In the event that an AI agent engages in illegal activity, it is unclear who bears the responsibility — the user, the provider, or the content creator.

See also: ChatGPT data leak and new cyberattacks

Google DeepMind: How do hackers breach AI Agents?

The future of trust in artificial intelligence

DeepMind's study comes to a key conclusion: the internet is no longer designed just for humans, but also for machines. This is fundamentally changing the security landscape.

The key question now is not just what information is available, but what information will AI systems trust. In a world where AI agents make decisions, data reliability becomes more critical than ever.

📧
Subscribe to the SecNews Newsletter

The most important Security & Technology news in your Inbox.

Digital Fortress
Digital Fortresshttps://www.secnews.gr/politiki-syntaxis/
Member of the SecNews Editorial Team. Covers software vulnerabilities, data breaches, cyberattacks and technology developments. All articles follow the SecNews Editorial Policy.

SEARCH

FOLLOW US

📧
Newsletter SecNews
The most important Security & Technology news in your inbox.

LIVE NEWS