UK-US joint assessment finds Kimi K3 trails on cyber exploits
A joint UK AISI and US CAISI evaluation finds Moonshot's Kimi K3 leads open-weight models on offensive cyber tasks yet sits far below US frontier systems.
The UK AI Security Institute (AISI) and the US Center for AI Standards and Innovation (CAISI) published a joint preliminary assessment of Moonshot AI's Kimi K3, concluding that the much-hyped Chinese open-weight model performs significantly below recent US frontier models on offensive cyber capabilities — while still setting a new high-water mark among open-weight systems.
The numbers
On ExploitBench, a Carnegie Mellon-developed benchmark built around 41 real vulnerabilities in Chrome's V8 engine, Kimi K3 scored about 32 percent on exploit-development progress, versus roughly 50 percent for leading US models — and produced zero complete arbitrary-code-execution exploits, where top US models solved 20 of the 41 tasks. On "The Last Ones," a simulated 32-step corporate network attack, K3 reached step 17 on average against 28.5 for the most capable US systems. K3 nonetheless beat Zhipu's GLM-5.2, the previous best open-weight model tested, which scored 24 percent on ExploitBench. The institutes also noted K3's built-in safeguards did not stop it from attempting exploit development.
The distillation subtext
The Decoder's analysis of the report adds a pointed interpretation: the gap is consistent with allegations that Chinese labs bootstrap models by distilling US frontier outputs. The evaluators tested US models with system-level safeguards disabled, exposing cyber capabilities that are nearly impossible to reach through public interfaces — and therefore unavailable to any lab distilling from those interfaces. That reading lands in a charged week, following US Treasury threats of sanctions over alleged distillation of Anthropic's Fable model.
Why it matters
This is the first joint UK-US government evaluation of a Chinese frontier model, and it turns a geopolitical shouting match into measurable claims: open-weight models are closing the general-capability gap but still lag badly where dangerous capability matters most. Expect both sides to weaponize the result — Washington as evidence that export controls and interface restrictions work, and open-source advocates as evidence that open weights are not the biosecurity-grade threat some policymakers describe.