The bad behavior is called reward hacking. Here's what you need to know. When two OpenAI AI agent models hacked the Hugging Face website in July, they weren't trying to make money or cause sabotage — they were simply looking for answers to a test question.
See also: Xbox Outage: Why weren't physical discs working?

According to an analysis by OpenAI, the models, which had been stripped of their typical security features for testing, decided to solve a cybersecurity exercise by breaching the isolated environment in which OpenAI had tried to confine them and hacking into the Hugging Face databases, where they believed the correct answer to the problem might be stored.
The Hugging Face incident has attracted intense attention in recent weeks. It’s a dramatic illustration of how good AI models have become at hacking: To break into Hugging Face’s databases, the models had to combine a number of previously unknown cybersecurity exploits. But it’s perhaps even more striking as an example of how and why AI systems lie and deceive. As the models become more powerful, the consequences could become much more serious.
Researchers have long known that AIs tend to take creative approaches to achieving the goals set for them. Back in 2016, Anthropic co-founders Dario Amodei and Jack Clark, who were then working at OpenAI, published a blog post about an AI agent they had trained to play a Flash boat racing game called Coast Runners.
Instead of driving through the race to the finish line, as the researchers had predicted, the agent found a corner of the course where it could spin around collecting power-ups, thereby maximizing its score. The Coast Runners story quickly became one of the most famous examples of reward hacking, a phenomenon in which AI agents complete tasks or earn high scores using unintended strategies.
See also: Why are CD sales suddenly rising again?

Historically, researchers have discussed reward hacking almost exclusively in the context of reinforcement learning, a common AI training regime. Like dog training, reinforcement learning involves providing a reward to a subject when it achieves a goal. Rewards reinforce the behaviors that led to that achievement. In the case of AI training, rewards are purely mathematical, but in essence they are the same as a dog treat: After receiving a reward, the agent is more likely to repeat whatever action produced it.
It can be difficult to write good rules for when and when not to give an agent a reward. In the case of Coast Runners, the agent was rewarded based on his score in the game and found a shortcut to achieving the highest possible score by spinning in circles for power-ups. Once he happened upon this strategy and received a reward for it, the strategy was reinforced and the agent abandoned the game altogether.
The solution was to modify the rewards by giving the agent fewer points for achieving power-ups and more for completing the route.
With today’s sophisticated LLM-based agents, determining when and when not to give a reward can be much more complex. If an AI system is asked to solve a coding problem, it might work hard to find the solution—the kind of behavior that AI companies want to reinforce. But it could also modify the code that evaluates whether the problem has been solved, search for the solution online, or cheat in other ways.
These are behaviors that AI companies want to eliminate from their models, but if the model cheats convincingly enough, it will be rewarded instead and the behavior will be reinforced.
See also: Why Apple is reportedly skipping the M6 Pro and M6 Max chips

Anthropic said it has detected some instances of cheating in its models during training, suggesting that other forms of cheating may be going undetected. If so, the models could be trained to behave badly.
🔒 Protect your privacy with Proton VPN
Swiss VPN from the creators of Proton Mail — strict no-logs policy, strong encryption, and built-in NetShield that blocks ads, trackers, & malware.
- ✔ No-logs, based in Switzerland (except 14-Eyes)
- ✔ NetShield: blocks ads, trackers & malicious domains
- ✔ Covers all devices — free version available
The link is an affiliate link — SecNews may receive a commission at no additional cost to you. It does not affect the independence of our article writing.
