Updated
Updated · NIST · Jul 23
Kimi K3 Tops GLM-5.2 at 32% but Misses 0 of 41 ACE Exploits
Updated
Updated · NIST · Jul 23

Kimi K3 Tops GLM-5.2 at 32% but Misses 0 of 41 ACE Exploits

3 articles · Updated · NIST · Jul 23

Summary

  • UK AISI and CAISI found Moonshot AI’s Kimi K3 beat GLM-5.2 on ExploitBench, scoring 32% versus 24% for the strongest open-weight model tested in June.
  • That gain did not extend to the highest-severity exploit stage: Kimi K3 achieved arbitrary code execution on 0 of 41 tasks, while the most cyber-capable models averaged 20 of 41.
  • On the 32-step “The Last Ones” cyber range, Kimi K3 reached step 17 on average, ahead of GLM-5.2’s 11 but well below leading U.S. models’ 28.5.
  • Kimi K3 completed the full attack path in 1 of 10 attempts within a 100 million-token limit, indicating it can autonomously attack small, weakly defended networks given initial access.
  • The agencies called the results preliminary because Kimi K3’s aggregate cyber score was estimated from a single 41-task benchmark; the model was released July 16 and is slated for open-weight release by July 27.

Insights

With Kimi K3's open-weight release days away, how will its autonomous attack capabilities transform the underground cybercrime economy?
How will the removal of safety guardrails from Kimi K3 impact enterprise security once it becomes freely available on July 27?