⚡ AI Focus Bulletin
SafetyResearchPolicy

FAR.AI leaderboard finds 448 universal jailbreaks in Grok 4.5

An independent audit of four frontier models found hundreds of universal jailbreaks in Grok 4.5 and Gemini 3.1 Pro, and none in Claude Fable 5 or GPT-5.6 Sol.

FAR.AI published an AI Security Leaderboard on Wednesday that ranks how well the safeguards on frontier models hold up against the attacks most likely to be aimed at them, and reported a gap between labs that is wider than an order of magnitude.

The nonprofit research group assembled a taxonomy of more than 60 publicly documented jailbreak techniques, built tooling to recombine them automatically, and then ran 1,000 randomly assembled attacks plus 500 expert-guided attacks against each model across five high-risk domains: chemical, biological, radiological and nuclear, explosives, and cybersecurity. An attack counted as a "universal jailbreak" only if it succeeded on more than three-quarters of the harmful requests in a domain.

Under those identical conditions, xAI's Grok 4.5 yielded 448 distinct universal jailbreaks and Google's Gemini 3.1 Pro yielded 249. Anthropic's Claude Fable 5 and OpenAI's GPT-5.6 Sol yielded none, in any domain, under any search strategy. The cost asymmetry is the sharper number: a working universal jailbreak cost roughly $58 of compute to discover on Grok 4.5 and about $278 on Gemini 3.1 Pro, while FAR.AI spent more than $14,200 per model on the two clean systems without finding one. Adam Gleave, FAR.AI's chief executive, said the distance between the models is far larger than most observers assume.

FAR.AI framed the result as a floor rather than a clearance: a model with no universal jailbreak in this evaluation is not certified secure, because the team tested a representative rather than exhaustive set of attack surfaces. It also stressed that every weakness it found belongs to a publicly described class of attack with publicly described defenses, making the outcome a product decision rather than an unsolved research problem.

Why it matters: Regulators, enterprise buyers and cloud distributors have had no common yardstick for comparing model safeguards, and vendors have graded their own homework. A reproducible, cross-lab ranking with a dollar cost attached converts safety claims into a procurement variable — and puts specific numbers behind the argument that the CBRN and cyber guardrails on two of the four leading models are, today, cheap to defeat.

Sources