Microsoft has announced the development of a lightweight scanner, designed to detect backdoors in open-source large language models (LLMs) and enhance trust in artificial intelligence (AI) systems.

Microsoft's AI Security team said the scanner uses three observable signals to reliably detect the presence of backdoors while maintaining a low rate of false positives.
According to Blake Bullwinkel and Giorgio Severi, these signatures are based on how trigger inputs measurably affect the internal behavior of a model, providing a powerful and meaningful basis for detection.
See also: DEAD#VAX campaign deploys AsyncRAT via Phishing VHD files
LLMs: What risks are there?
LLMs can be vulnerable to two types of tampering: the model weights, which are the configurable parameters that support the decision-making logic, and the code.
Another form of attack is model poisoning, where a threat actor embeds hidden behavior directly into the model weights during training, leading to unintended actions when specific triggers are detected. These backdoored models remain inactive until the trigger is activated. This hidden attack allows a model to appear normal in most situations, while responding differently under narrowly defined trigger conditions.
See also: 1.5 million AI agents are not adequately monitored
Vulnerable LLMs: How Microsoft's scanner helps
Microsoft's study identified three practical signals that can indicate a poisoned AI model:
1. If a prompt contains a trigger phrase, poisoned models exhibit a characteristic “double triangle” attention pattern, which forces the model to focus on the trigger in isolation, as well as dramatically collapsing the “randomness” of the model’s output.
2. Backdoored models tend to leak their own poisoning data, including triggers, through memorization (instead of training data).
3. A backdoor introduced into a model can be activated by multiple “fuzzy” triggers, which are partial or approximate variations.

Microsoft explained that their approach is based on two key findings: first, sleeper agents tend to memorize poisoning data, allowing backdoor examples to be leaked via memory extraction.
Second, poisoned LLMs show characteristic patterns in their output distributions and attention heads when there are backdoor triggers in the input.
These indicators can be used to scan models at scale for built-in backdoors. This scanning methodology does not require additional model training or prior knowledge of backdoor behavior and is effective on common GPT-type models.
🔒 Protect your privacy with Proton VPN
Swiss VPN from the creators of Proton Mail — strict no-logs policy, strong encryption, and built-in NetShield that blocks ads, trackers, & malware.
- ✔ No-logs, based in Switzerland (except 14-Eyes)
- ✔ NetShield: blocks ads, trackers & malicious domains
- ✔ Covers all devices — free version available
The link is an affiliate link — SecNews may receive a commission at no additional cost to you. It does not affect the independence of our article writing.
See also: Chinese hackers Amaranth-Dragon exploit WinRAR vulnerability
Microsoft's scanner extracts memorized content from the model, analyzes it to isolate significant substrings, and formulates the three signatures as loss functions to score suspicious substrings, returning a ranking of candidate triggers.

However, the scanner has limitations. It does not work on proprietary models, due to the need to access the model files. It is also more effective on trigger-based backdoors that produce deterministic outputs, and cannot detect all types of backdoor behavior.
The researchers consider this work an important step towards practical backdoor detection and emphasize that continued progress relies on shared learning and collaboration within the AI security community.
