HomeSecurityHackers use Gemini tools against him

Hackers use Gemini's tools against him

They say it takes a thief to catch a thief, and perhaps the same is true when it comes to hacking LLMs. Ethical hackers have discovered a way to make Google's Gemini AI models more vulnerable to hacking — and they did it using Gemini's own tools.

See also: Gemini's Astra is rolling out to some Android users

Gemini hacker

The technique was developed by a team from UC San Diego and the University of Wisconsin , as reported by Ars Technica . It’s called “ Fun-Tuning ,” and it significantly increases the success rate of direct injection attacks, where hidden instructions are embedded in text that an AI model reads . These attacks can cause the model to leak information , give incorrect answers, or perform other unintended actions.

What makes the method interesting is that the hackers are using Gemini’s fine-tuning feature, which is usually intended to help businesses train AI on custom datasets. Instead, the researchers used it to automatically test and improve the injections. It’s like teaching Gemini how to fool itself.

See also: Gemini may soon respond to your video uploads

Fun-Tuning works by creating strange prefixes and suffixes that are added to an otherwise ineffective prompt. These additions “boost” the prompt and make it much more likely to succeed. In one case, a prompt that failed on its own was made effective by wrapping it in suffixes like “wandel !!! !!!” and “formatted ! ASAP !“

Hackers use Gemini's tools against him

In testing, hackers had a 65% on Gemini 1.5 Flash and an 82% on the older Gemini 1.0 Pro — more than double the baseline success rates without Fun-Tuning. The attacks also transferred well between models, meaning that an injection that worked on one version often worked on others.

The vulnerability stems from the way fine-tuning works. During the prompt, Gemini provides feedback in the form of a “loss” score, which is a number that reflects how far the model’s answer is from the desired outcome. Attackers can exploit this feedback to fine-tune their prompts until the system finds a successful outcome.

See also: How the new Gemini model uses your search history

The researchers believe the problem may be difficult to fix, as the feedback that Fun-Tuning allows is a key part of how it works. In other words, making Fun-Tuning less effective risks reducing its usefulness overall.

Source: androidauthority

Selecting the team

🔒 Protect your privacy with Proton VPN

Swiss VPN from the creators of Proton Mail — strict no-logs policy, strong encryption, and built-in NetShield that blocks ads, trackers, & malware.

  • ✔ No-logs, based in Switzerland (except 14-Eyes)
  • ✔ NetShield: blocks ads, trackers & malicious domains
  • ✔ Covers all devices — free version available
Try Proton VPN for free — 30-day money-back guarantee →

The link is an affiliate link — SecNews may receive a commission at no additional cost to you. It does not affect the independence of our article writing.

📧
Subscribe to the SecNews Newsletter

The most important Security & Technology news in your Inbox.

Absentee Mia
Absentee Miahttps://www.secnews.gr/politiki-syntaxis/
Member of the Editorial Team of SecNews. He writes about cybersecurity, online fraud, privacy and technology. All articles follow the SecNews Editorial Policy.

SEARCH

FOLLOW US

📧
Newsletter SecNews
The most important Security & Technology news in your inbox.

LIVE NEWS