I recently got to watch what happens when you jailbreak some of the world’s most powerful artificial intelligence models. Don’t worry—this AI manipulation wasn’t used to hack anyone or build a nuclear bomb. I simply got to see firsthand how vulnerable some frontier models are to ditching their safety guardrails. FAR.AI, an AI safety nonprofit based in California, built a tool that takes a range of problematic prompts and generates more than a thousand different versions in an attempt to identify functioning jailbreaks. I saw some models generate a detailed plan for launching a cyberattack on a...
It’s Frighteningly Easy to Jailbreak Some Frontier AI Models