Kimi K3 Tops GLM-5.2 at 32% but Misses 0 of 41 ACE Exploits
Updated
Updated · NIST · Jul 23
Kimi K3 Tops GLM-5.2 at 32% but Misses 0 of 41 ACE Exploits
3 articles · Updated · NIST · Jul 23
Summary
UK AISI and CAISI found Moonshot AI’s Kimi K3 beat GLM-5.2 on ExploitBench, scoring 32% versus 24% for the strongest open-weight model tested in June.
That gain did not extend to the highest-severity exploit stage: Kimi K3 achieved arbitrary code execution on 0 of 41 tasks, while the most cyber-capable models averaged 20 of 41.
On the 32-step “The Last Ones” cyber range, Kimi K3 reached step 17 on average, ahead of GLM-5.2’s 11 but well below leading U.S. models’ 28.5.
Kimi K3 completed the full attack path in 1 of 10 attempts within a 100 million-token limit, indicating it can autonomously attack small, weakly defended networks given initial access.
The agencies called the results preliminary because Kimi K3’s aggregate cyber score was estimated from a single 41-task benchmark; the model was released July 16 and is slated for open-weight release by July 27.