Researchers have developed an automated data poisoning tool that they say can render stolen, high-value data used in AI systems useless, a solution that CSOs may need to adopt to protect sophisticated large language models (LLMs). The technique, created by researchers from universities in China and Singapore, involves injecting convincing but false data into what is known as a knowledge graph (KG) created by an AI operator.
See also: Traditional security frameworks leave organizations exposed to AI attacks

A knowledge graph contains the proprietary data used by LLM. The introduction of poisoned or corrupted data into a data system to protect against theft is not new.
What's new about this tool – called AURA (Active Utility Reduction via Adulteration) – is that authorized users have a secret key that filters out fake data so that the LLM's answer to a query is useful. However, if the knowledge graph is stolen, it is useless to the attacker unless they know the key, as the adulterated data will be retrieved as context, causing degradation in the LLM's logic and leading to factually incorrect answers.
The researchers say that AURA reduces the performance of unauthorized systems to just 5.3%, while maintaining 100% fidelity for authorized users, with “negligible overhead,” defined as a maximum query latency increase of less than 14%. They also say that AURA is resistant to various sanitization attempts by an attacker, retaining 80.2% of the corrupted data entered for defense, and the fake data it creates is difficult to detect.
Why is all this important? Because KGs often contain an organization's highly sensitive intellectual property (IP), they are a valuable target. However, the proposal has been met with skepticism by one expert and caution by another.
See also: The dark side of algorithms: When AI becomes the target of cyberattacks

“Data poisoning has never really worked well,” said Bruce Schneier, chief security architect at Inrupt Inc., and a fellow and lecturer at Harvard’s Kennedy School. “Honeypots, no better. That’s a clever idea, but I don’t see it as anything more than a security adjunct.” Joseph Steinberg, a U.S.-based cybersecurity and AI consultant, disagreed, saying, “generally this could work for all kinds of AI and non-AI systems. This is not a new concept.”
Graph of Knowledge 101: LLMs use a technique called RAG to search for information based on a user query and provide the results as additional reference for generating AI system responses. In 2024, Microsoft introduced GraphRAG to help LLMs answer queries that require information beyond the data they are trained on.
GraphRAG uses knowledge graphs generated by LLM to improve performance and reduce the likelihood of response illusions when performing discovery on private datasets such as a company’s proprietary research, business documents, or communications. The proprietary knowledge graphs within GraphRAGs make them “a prime target for IP theft,” like any other proprietary data.
See also: AI red teaming: Testing the security of the models themselves

AI is moving faster than AI security. As AI usage spreads, CSOs need to remember that artificial intelligence and everything it takes to make it work also makes it much harder to recover from bad data introduced into a system, Steinberg noted.
