Jailbreaking AI-Powered Robots is Way Too Easy: Study

Jailbreaking an AI Powered Robot Illustration

It’s terrifyingly easy to jailbreak an AI-powered robot.

That’s according to researchers from Penn Engineering who’ve developed an algorithm, called RoboPAIR, that effortlessly bypasses robotic safety systems. In their tests, they achieved a 100 percent success rate by using prompts specifically designed to confuse large language models (LLMs).

The robots they tested included the Unitree Go2, using GPT-3.5, the Clearpath Jackal, which was equipped with GPT-4o, and NVIDIA’s Dolphin self-driving simulator, powered by a fine-tuned OpenFlamingo model. Each of these AI systems was vulnerable to prompt-based manipulation, which allowed the researchers to override built-in safety protocols.

They were able to bypass the safety systems of the Unitree Go2 quadrupedal robot and the Jackal unmanned ground vehicle, showing that these robots could be manipulated into unsafe behaviors. Most strikingly, they got NVIDIA’s Dolphin self-driving simulator to speed recklessly through crosswalks, ignoring critical safety rules.

In the end, the researchers concluded that AI-powered robots are not yet ready for real-world deployment. Their vulnerabilities make it too easy for bad actors to manipulate them, potentially turning once-safe systems into hazardous threats.