Researchers expose a worryingly simple trick to make AI bots go rogue and skip safety
This new study reveals how patient, step-by-step manipulation can trick AI agents into ignoring their own safety rules.
Aerps.com / Unsplash
If you ask an AI agent to hack an account, it will most certainly refuse, but researchers at EPFL just proved there is an easier way in, and it involves patience rather than technical skill. Their new study shows that breaking a harmful goal into small, harmless-sounding requests can trick AI agents into completing tasks they would normally reject outright (via TechXplore).
It echoes the recent ‘Bioshocking’ exploit in which AI browsers were manipulated into treating credential theft as part of a harmless game.
How researchers exposed this weakness
The team built an automated testing tool called STING, short for Sequential Testing of Illicit N-step Goal execution, designed to mimic how a real attacker would actually operate. Instead of stating a harmful goal directly, STING plans ahead and breaks that goal into a sequence of smaller, seemingly innocent steps that build toward it over multiple conversation turns.
Helpful to a Fault: Measuring Illicit Assistance in Multi-Turn, Multilingual LLM Agents Source
Researchers tested this approach across 176 harmful scenarios against leading AI models, including ChatGPT, Gemini, and Claude. Each AI agent was tested as a tool-using agent, capable of browsing the web, sending emails, and completing multistep tasks.
Gradual, multistep manipulation succeeded far more often than blunt, single-prompt attempts. In some cases, models were twice as likely to complete a harmful task once the request was broken down into smaller steps. That finding tracks with separate research showing that even average users can talk their way past AI safety guardrails using nothing more than carefully worded prompts.
Why does this matter?
The concerns raised by researchers is not a hypothetical risk. Meta admitted in June that attackers used simple social engineering, not malware or hacking tools, to trick its AI support assistant into granting unauthorized access to Instagram accounts.
Unsplash
The researchers also expected attacks to be more effective in languages with less available training data. However, they found that completion rates stayed roughly consistent across all seven languages tested. They found one exception, though: switching languages midway through a multi-step attack made success rates jump significantly.
Lead researcher Ayush Kumar Tarun argues that safety testing needs to happen much earlier, built into an agent’s design from the start. Bolting it on after something goes wrong is no longer good enough, especially as these systems keep gaining more real-world capabilities.

Manisha Priyadarshini is a tech and entertainment writer with over nine years of editorial experience.
[Update] Googlebook finally gets a launch date, and it’s sooner than you would expect
Googlebook gets a September 15 showcase with Gemini, Android and new laptops
Update: Google clarified information surrounding the event. Here is the statement: "Nothing will be announced on September 15; that is just the date of the embargoed event."
The original story is as follows.
Intel’s next laptop CPUs could get a huge cache upgrade with Razor Lake
Razor Lake is rumored to bring bLLC to mobile chips, while an earlier N2X claim has already been revised to N2P V2

Intel’s future laptop chips could be in line for a much bigger cache. Leaker Jaykihn says Razor Lake will include mobile SKUs with bLLC, or Big Last Level Cache, potentially extending Intel’s large-cache strategy to notebooks.
There’s already reason to be cautious with the details. Jaykihn initially identified TSMC’s N2X process for Razor Lake before correcting the claim to N2P V2. Intel hasn’t announced Razor Lake or confirmed either specification, so this remains a very early look at the generation.
Windows 11’s latest update has become a gamble for gamers
KB5121003 is crashing The Finals for some players, while Microsoft still lists no known issues with the update

Windows 11’s August update is becoming a risky install for some gamers. Windows 11 update problems are hardly new, but KB5121003 has now been linked to crashes in The Finals, and developer Embark has already published troubleshooting steps for players who run into them.
Embark traced the problem to inpoutx64.sys, a third-party driver that can conflict with the update. Its workaround removes the troublesome driver while leaving KB5121003 itself installed, which is a much better option than rolling back the whole update.
KickT