Moonshot released Kimi K3’s 2.8 trillion-parameter weights on 26 July, two days ago, and one day earlier than expected. The UK AISI / CAISI assessment published three days before that found K3 scores 32% on ExploitBench against GLM-5.2’s 24%, reaches arbitrary code execution on 0 of 41 targets where leading US models average 20, and averages step 17 of 32 on a simulated intrusion range against the frontier’s 28.5. Three caveats sit under those numbers: the US models were tested with safeguards switched off, Kimi K3 got a narrower benchmark set because of its hosting, and its score carries a wider confidence interval than any other model’s. The capability finding that survives all three is the one AISI ran with safeguards fully on: 1 in 10 attempts completed the whole 32-step range. Treat the ranking as soft and the 1-in-10 as the number to plan around.
I wrote up this assessment four days ago, the day after UK AISI and the US Center for AI Standards and Innovation published it. Two things have changed since. Moonshot put the weights online, and someone finally did the arithmetic on how the comparison was built.
The comparison has safeguards off on one side and fewer tests on the other
The headline is that Kimi K3 “trails leading US frontier closed weight models on cyber capability.” AISI is upfront that the US closed-weight models were “evaluated with system-level safeguards disabled to reduce refusals and enable measurement of maximal capabilities.” That is the correct way to measure raw capability. It also means the frontier’s score is a ceiling, and the number you would get from the public API is lower.
The other side of the comparison is trimmed too. NIST’s mirror of the release says that “due to the specifics of Kimi K3’s hosting setup, UK AISI / CAISI ran a selective set of cyber evaluations.” At the time of testing, K3 was API-only, having launched 12 days ago with no downloadable weights.
XenoSpectrum put the objection plainly three days ago: “Concluding that ‘the gap is stark’ by placing a US model with safety measures stripped away side by side with a Kimi K3 tested under a narrowed evaluation scope surely calls for a further layer of explanation.” My original post flagged the safeguards asymmetry. It missed that the scope was narrowed on the other side as well, which makes the ranking softer than a single percentage gap implies.
Third asterisk, from the same NIST mirror: Kimi K3’s score “has a larger confidence interval than other models because it was estimated from a single benchmark.” A one-benchmark estimate with an acknowledged wide interval does not support a precise league-table position.
The report never says which US models it beat Kimi K3 with
The Hacker News thread, 131 points and 54 comments, converged on the same gap. Neither the AISI post nor the NIST mirror defines “top US models,” and without the set you cannot tell whether the frontier figure is a mean, a median, or the best single run. One commenter says the full NIST report names them as OpenAI’s GPT-5.5 and Anthropic’s Mythos Preview. I have not verified that against the full report, so treat it as unconfirmed.
This matters more than benchmark pedantry usually does, since the whole policy value of the number is comparative. “Trails the frontier by X” is only meaningful with a defined frontier.
What survives the caveats
Strip out everything that depends on the cross-model comparison and the load-bearing findings are still there, because they are measurements of Kimi K3 alone.
- Its safeguards “did not prevent it from attempting cyber exploit development or offensive cyber operations.” No asterisk needed on that one, since nothing was disabled to produce it.
- In 1 of 10 attempts it completed the full 32-step simulated network intrusion. My earlier post describes the range in more detail, drawn from the AISI page: four subnets, roughly 20 hosts, about 20 hours of work for a human expert.
- AISI’s own read: K3 “is capable of autonomously attacking small, weakly defended and vulnerable enterprise systems, when directed to do so and given initial network access.”
The range has no live defenders, no alerting penalty and a deliberate attack path. A 1-in-10 solve against a soft target is a weak autonomous attacker. It is also, per AISI, no longer rare: “Solves of TLO are no longer exclusive to a small set of models.”
Downloadable as of two days ago
Moonshot released the weights on 26 July, with Together AI and Modal announcing day-0 hosting. My earlier post said the release was due 27 July, which was the date circulating before it happened; the reported date is a day earlier. I have not found a primary Moonshot announcement confirming whether that shipped the full weights or a staged subset, so the one-day slip is reported, not nailed down.
Either way the safeguards question closes. Whatever refusal behaviour remains in K3 can be fine-tuned out by anyone with the weights, and the finding that its safeguards never fired during exploit development means there was little to remove.
The timing is worth weighing
XenoSpectrum’s other point is political. The assessment landed 23 July, one day after OSTP director Michael Kratsios publicly accused Moonshot of distilling Anthropic’s Fable to build K3, an accusation I covered when it was made and found thin on evidence.
Convenient timing is not falsification. AISI’s methodology is documented, its caveats are in its own text, and this is the second cyber gap measurement it has published this month after the 17 July open-weight analysis that put open models four to seven months behind. The institute was working this beat before the distillation row started.
But a number produced by two governments, about a Chinese lab their administrations are actively trying to restrict, in the same week as a distillation accusation and sanctions chatter, will get quoted stripped of every caveat in its own footnotes. The wide confidence interval will not survive contact with a press release. The 1-in-10 full-range solve, measured on a model whose safeguards did nothing and whose weights are now public, is the part that deserved the headline.
Sources
- AISI / CAISI, UK AISI / CAISI Preliminary Assessment of Kimi K3’s Cyber Capabilities (23 July 2026); NIST mirror
- XenoSpectrum, Why Is Kimi K3’s Cyberattack Capability Less Than Half of the US Level? (25 July 2026)
- Simon Willison, Kimi K3 (16 July 2026)
- AISI, How far behind the frontier are leading open-weight models on cyber? (17 July 2026)
Coverage
- Hacker News: Preliminary Assessment of Kimi K3’s Cyber Capabilities (131 points, 54 comments)