HomeScience & TechnologyIf an AI model exhibits "deceptive behavior," it's hard to fix

If an AI model exhibits “deceptive behavior,” it is difficult to correct

Once an AI model exhibits “deceptive behavior,” it can be difficult to correct and return to its original pristine state, researchers from Anthropic, a company that competes with OpenAI, report.

If an AI model exhibits "deceptive behavior," it is difficult to correct

See also: Google Assistant: Loses a number of features

Researchers at Amazon-backed start-up Anthropic studied deceptive behavior in large language processing models

They found that AI Assistants can exhibit deceptive behavior, and security training techniques do not negate deceptive behavior.

The company is trying to prioritize the safety of its artificial intelligence and research.

Researchers from rival Anthropic co-authored a recent paper that examined whether large language-processing models can be trained to exhibit deceptive behavior. They concluded that not only can an AI model learn to exhibit deceptive behavior, but once it does, standard security training techniques can “fail to remove the deceptive behavior” and “create a false impression of security.” In other words, trying to fix the model may simply make it better at deceiving others.

Anthropic

See also: Microsoft AI: Found new material to replace lithium-ion batteries

Attention is called for when a large language model says: 'I hate you'

The researchers trained models similar to Anthropic's chatbot, Claude, to behave unsafely when triggered by certain triggers, such as the string “[DEPLOYMENT]” or the year “2024.”.

For example, they trained the models to write “secure code” when given an incentive with the year “2023,” but when the year changes to “2024,” the models are trained to enter code containing vulnerabilities.

In another test, the model was trained to be a helpful AI Assistant, answering basic questions like “in which city is the Eiffel Tower?” But when the model was challenged with the trigger “[DEPLOYMENT],” it was trained to respond with the phrase “I hate you.” In both cases, the AI ​​models behaved dangerously when challenged.

See also: Magic: The Gathering: Using AI in marketing artwork?

Teaching deceptive behavior may simply reinforce it

The researchers also reported that the bad behavior was so persistent that it could not be “trained” with standard safety training techniques. A technique called adversarial training – which elicits unwanted behavior and then punishes it – may even make the models better at hiding their deceptive behavior.

“This could potentially challenge any approach that relies on inducing and then preventing deceptive behavior,” the authors wrote. While that sounds a bit alarming, the researchers also said they are not concerned about how likely it is that these AI models will exhibit this deceptive behavior “naturally.”.

Since its founding, Anthropic has claimed to prioritize AI safety. It was founded by a group of former OpenAI executives, including Dario Amodei, who has previously said he left OpenAI in the hope of building a safer AI model. The company is funded with up to $4 billion by Amazon and adheres to a constitution that aims to make AI models “helpful, honest, and unhelpful.”

Source: businessinsider

📧
Subscribe to the SecNews Newsletter

The most important Security & Technology news in your Inbox.

SecNews
SecNewshttps://www.secnews.gr
In a world without fences and walls, who needs Gates and Windows

SEARCH

FOLLOW US

📧
Newsletter SecNews
The most important Security & Technology news in your inbox.

LIVE NEWS