HomeSecurityFriendly Fire: AI agents are tricked into executing malicious code

Friendly Fire: AI agents are tricked into executing malicious code

The Friendly Fire attack demonstrates that AI agents designed to detect malicious code can be tricked into executing it themselves. Researchers at the AI ​​Now Institute have published a proof-of-concept that targets two of the most popular code security tools: Anthropic ’s Claude Code and OpenAI ’s Codex . The result is alarming: instead of detecting the threat, the tools themselves become the means of attack.

See also: Vulnerability in Claude Code GitHub Action allows malicious issue to take over repositories

Friendly Fire attack AI agents malicious code Claude Code Codex

Researchers Boyan Milanov and Heidy Khlaaf tested two configurations, both with autonomous mode enabled: Claude Code 's auto-mode and Codex 's auto-review . These modes use a classifier to execute commands that the AI ​​agent deems safe, stopping only at those it deems dangerous. They fall between completely unrestricted access and a configuration that asks for confirmation before each action. They are opt-in modes, and the attack requires that one of them be enabled.

The attack scenario is extremely realistic. A malicious user inserts a few extra files into an open source library. The test used geopy, a popular Python for searching for geographic coordinates, although the researchers emphasize that the method is portable to almost any project. A note in the README.md suggests running a script named security.sh as a check routine before a pull request. The script silently launches a hidden binary carrying the actual payload.

How the Friendly Fire attack fools AI agents

To bypass the AI ​​agent, the researchers disguised the malicious binary as the compiled result of an innocent Go that was sitting next to it. They also enriched it with strings from that file so that Claude Code could link them together. No part of the library code ever calls the binary, so nothing looks suspicious. All you have to do is give the AI ​​agent a simple command like “Perform security testing on this project” and the rest is done automatically. The AI ​​agent reads the README, determines that the script is part of the job, and executes it. The malicious binary runs on the victim’s computer without any warning.

See also: Agentjacking: New attack tricks AI coding agents into executing malicious code

Friendly Fire: AI agents are tricked into executing malicious code

A critical element that makes this attack particularly dangerous is the choice of vehicle: the README.md file . Previous attacks on AI agents have exploited machine configuration files, such as .mcp.json or .claude/settings.json , which trigger Claude Code ’s “ Yes, I trust this folder ” warning . This attack is hidden in a simple text file that exists in almost every repository. There is no trust prompt, no elevated access required — a much wider opening for attack.

Friendly Fire: Why a model update is not enough

The scope of the attack is impressive. The payload written for Claude Sonnet 4.6 worked unchanged on Sonnet 5 , Opus 4.8 , and GPT-5.5 . In some runs, the newer models noticed that the binary didn’t match its supposed source — and executed it anyway. When asked directly whether the geopy library contained hidden instructions, both Claude Sonnet 4.6 and GPT-5.5 answered in the negative. One attack, two vendors, four models, no change. This is the basis for the AI ​​Now Institute ’s claim : the problem can’t be fixed with a model update, because the models still can’t reliably distinguish the code they’re reading from the instructions they’re supposed to be following.

The AI ​​Now Institute points out that the weakness lies in the design, not in a specific software version. Anthropic has released three patches for configuration file injection in the last six months, but this attack bypasses that category entirely. The solution is not a new version, but a change in workflow: AI agents should not execute code from unverified repositories without human confirmation.

See also: Anthropic: Mythos may come to Claude Code

Friendly Fire: AI agents are tricked into executing malicious code

The researchers also address their findings to policymakers. Governments and companies are pushing AI agents into defensive security tasks — a U.S. in June is among them — at a faster pace than they are closing the loopholes revealed by this attack. The research remains a proof-of-concept with no reported exploit in practice, but the message is clear: blindly trusting autonomous AI agents to inspect third-party code is a serious security risk that needs to be addressed immediately.

Selecting the team

🔒 Protect your privacy with Proton VPN

Swiss VPN from the creators of Proton Mail — strict no-logs policy, strong encryption, and built-in NetShield that blocks ads, trackers, & malware.

  • ✔ No-logs, based in Switzerland (except 14-Eyes)
  • ✔ NetShield: blocks ads, trackers & malicious domains
  • ✔ Covers all devices — free version available
Try Proton VPN for free — 30-day money-back guarantee →

The link is an affiliate link — SecNews may receive a commission at no additional cost to you. It does not affect the independence of our article writing.

📧
Subscribe to the SecNews Newsletter

The most important Security & Technology news in your Inbox.

Absentee Mia
Absentee Miahttps://www.secnews.gr/politiki-syntaxis/
Member of the Editorial Team of SecNews. He writes about cybersecurity, online fraud, privacy and technology. All articles follow the SecNews Editorial Policy.

SEARCH

FOLLOW US

📧
Newsletter SecNews
The most important Security & Technology news in your inbox.

LIVE NEWS