In a nondescript evaluation suite, researchers fed prompts into Moonshot AI's freshly released Kimi K3 and watched it navigate a simulated corporate network. The results, published on 23 July by the UK Artificial Intelligence Security Institute and its US counterpart, the Center for AI Standards and Innovation, offer a measured snapshot of where Chinese frontier efforts stand.
Kimi K3, launched on 16 July with an open-weight version due around 27 July, reached step 17 on average in a 32-step simulated corporate network attack known as The Last Ones. By contrast, the most cyber-capable US frontier models averaged 28.5 steps. The Chinese model managed to complete the full exercise in only one out of ten attempts within a 100-million-token limit.
On ExploitBench, a benchmark drawn from 41 recent vulnerabilities in the V8 engine, Kimi K3 scored 32 percent. It failed to achieve arbitrary code execution on any task. Leading US models succeeded on 20 out of 41 on average. The assessment also showed Kimi K3 outperforming GLM-5.2, the most capable open-weight model as of June, which scored 24 percent.
Capabilities that matter for national infrastructure
These numbers matter because the tests probe exactly the sort of autonomous offensive potential that could target real enterprise systems. Kimi K3's safeguards did not stop it from attempting exploit development or offensive operations during the evaluations. Yet it still fell significantly short of the best US systems, even when those were tested with safeguards removed to measure maximum potential.
The evaluations used a selective set of cyber tests, shaped by the specifics of Kimi K3's hosting setup. They represent early findings on a small set of public and private benchmarks. Such limits are inevitable at this stage, yet the pattern is consistent: Chinese developers are advancing, but established Western frontier models retain a clear, if shrinking, lead in the domains that matter most for security.
Why independent testing cannot be optional
The joint work underscores the practical value of rigorous, evidence-based assessments conducted by Western safety institutes. Rather than relying on developer claims or blanket assumptions about openness, these exercises measure what models can actually do. In an environment where commercial incentives push rapid release, such scrutiny provides governments and citizens with data grounded in reality.