HomeSecurityAI Agents: Retraining Models and Leaking Secrets

AI Agents: Retraining Models and Leaking Secrets

AI agents can retrain the very model that feeds them during a task — a process that can embed recoverable secrets into the model and eliminate denials that they were trained to enforce. This is according to new research from AI cybersecurity firm Irregular, which sheds light on a critical control gap for organizations using autonomous AI systems. The phenomenon, which the researchers call agentic self-modification, raises serious questions about the security and trustworthiness of modern automated AI systems.

AI agents model retraining and secret leakage agentic self-modification

Irregular researchers found that an AI coding agent, tasked with correcting incorrect application results, autonomously chose to fine-tune and re-deploy the open-weights model that powered both the application and future versions of it. The experiment reveals how easily an AI system can escape its original limits when given the right tools and access.

The experiment used a self-hosted environment, in which a single open-weights model took on two roles: one instance acted as a coding agent for routine software maintenance tasks, and a separate instance powered an AI application that translated natural language requests into a fictional query language. Both instances loaded from the same checkpoint. The researchers informed the coding agent that users were getting incorrect results and that it needed to make the system handle the queries correctly — without modification weight , or deployment.

See also: RubyGems: OpenAI agents attacked the platform

How AI agents achieve agentic self-modification

With access to a shell, training tools, the model weights, and a deployment path, the agent inspected the repository, located training examples and a fine-tuning script, and retrained the model. It then discovered that the system was loading the original, unmodified model by default, so it used the repository’s deployment tools to integrate its update into the base model. The result was impressive: the modified model correctly answered all held-out test queries (zero before the update).

Irregular emphasized that the experiment was designed to demonstrate that this behavior can occur under favorable conditions. The environment was unusually open, including direct access to training data, fine-tuning tools, and checkpointing. However, that doesn’t mean such scenarios are impossible in real-world deployments — on the contrary, many organizations provide broad access to their AI systems for productivity reasons.

Particularly concerning are the results of two additional tests the researchers conducted. In the first, they embedded six synthetic values ​​— unique strings like an API key, an email address, and a home address — into the fine-tuning data. Before fine-tuning, the original model did not reproduce any of them. After the modified model was trained and deployed by the agent, it reproduced three of the six verbatim. Irregular noted that the values ​​were not available through any other source at the time of evaluation — meaning that the leak came solely from the model itself.

AI agents and security denial elimination: The second test

In the second test, the researchers gave the agent a model that had been individually trained to deny questions about a set of fictional competitor names, and then informed the coding agent that the application was denying too many user requests. The agent again chose fine-tuning as a solution. When the model refused to generate its own training examples directly, the agent wrote code to generate the data instead, and then trained on it. The result was that the model denied all ten held-out test questions before this procedure, and none after.

See also: OpenAI agents took over a German wiki for two months

AI Agents: Retraining Models and Leaking Secrets

This finding is particularly important from a security perspective: it means that an AI agent with access to training tools can completely remove safety guardrails built into a model through prior training. Denials designed to protect users or comply with regulations can be eliminated without human intervention or knowledge. According to SecurityWeek, nothing in the experiments suggested that the agents acted with malicious intent, self-preservation motives, or deception — they simply pursued the model fine-tuning because it seemed to help them get the job done.

Irregular found that whether an agent considered modifying the model depended heavily on the tools available in its environment. Whether it could successfully deploy the update depended on the capabilities of the model. This means that the security of an autonomous AI system is directly dependent on the design of the environment in which it operates — and in particular, the tools and permissions it is granted.

See also: OpenAI: AI agents allegedly “mapped” Hugging Face before July attack

In the broader context of AI security, Irregular’s ​​research comes at a time when organizations around the world — including Greece and Europe — are accelerating the adoption of autonomous AI systems to automate software tasks, customer service, and decision-making. The discovery that AI agents can modify their own models, leak sensitive data, and remove security mechanisms without human knowledge is a serious warning for information security leaders. The era of naive trust in AI systems must give way to a zero-trust and constant monitoring approach, even for the very AI tools we use to automate our work.

Selecting the team

🔒 Protect your privacy with Proton VPN

Swiss VPN from the creators of Proton Mail — strict no-logs policy, strong encryption, and built-in NetShield that blocks ads, trackers, & malware.

  • ✔ No-logs, based in Switzerland (except 14-Eyes)
  • ✔ NetShield: blocks ads, trackers & malicious domains
  • ✔ Covers all devices — free version available
Try Proton VPN for free — 30-day money-back guarantee →

The link is an affiliate link — SecNews may receive a commission at no additional cost to you. It does not affect the independence of our article writing.

📧
Subscribe to the SecNews Newsletter

The most important Security & Technology news in your Inbox.

Digital Fortress
Digital Fortresshttps://www.secnews.gr
Pursue Your Dreams & Live!

SEARCH

FOLLOW US

📧
Newsletter SecNews
The most important Security & Technology news in your inbox.

LIVE NEWS