⚡ Uncle Cat AI Radar
SafetyResearch

FAR.AI leaderboard finds $58 universal jailbreak on Grok

The nonprofit's first AI Security Leaderboard reports a hundredfold spread in safeguard robustness, with two frontier models yielding none.

An apples-to-apples security ranking

FAR.AI launched its AI Security Leaderboard on August 4, an independent public ranking of how well frontier model safeguards hold up under the attacks most likely to be used against them. The nonprofit assembled a taxonomy of more than 60 publicly documented jailbreak techniques, built tooling to compose them automatically, then ran 1,000 randomly assembled attacks and 500 expert-guided attacks against each model across five risk domains: chemical, biological, radiological, nuclear, explosive and cyber.

The target was universal jailbreaks — reusable prompts that unlock whole categories of refused requests rather than a single lucky answer.

The spread

Results diverged sharply. xAI's Grok 4.5 yielded 448 distinct universal jailbreaks and Google DeepMind's Gemini 3.1 Pro 249. Automated random search alone accounted for 63 and 18 respectively; adding expert composition of techniques pushed those to 385 and 231, and the hand-built attacks were far more likely to work across three or more risk domains simultaneously. Claude Fable 5 and GPT-5.6 Sol yielded none — in any domain, under any search strategy.

FAR.AI converts that into cost. A working universal jailbreak ran roughly $58 to find on Grok 4.5 and roughly $278 on Gemini 3.1 Pro. On the two models that held, the search simply never succeeded, putting the floor above $14,200 and climbing. For open-weight models not yet on the board, the team says it has never needed more than a few hours, or roughly $10 to $50, to break one completely. Every vulnerability found belongs to a known attack class with defences that already exist and already run in deployed systems.

Why it matters

The finding is not that safeguards are weak but that they are wildly uneven — a hundredfold gap between shipping products, invisible to buyers and to the capability benchmarks the industry ranks on. Because the attacks are all previously documented, the gap reads as an engineering choice rather than an open research problem. It also sharpens the open-weight debate: if hosted models with active filtering fall for $58, the downloadable ones fall for pocket change.

Sources