Research: Kimi K3 network attack test score is 32.2%, trailing leading US models by more than 40 percentage points

robot
Abstract generation in progress

Deepchao TechFlow news: On July 24, a joint assessment by the UK AI Safety Institute and the US Center for AI Standards and Innovation found that Moonshot AI’s Kimi K3 can help with exploit development and simulated network attacks in the absence of substantial safeguards, but its overall cybersecurity capabilities still lag significantly behind leading US models.

On the ExploitBench exploitation benchmark, Kimi K3’s success rate is 32.2%, higher than China’s GLM-5.2, but far below the leading US model’s 76.2%, and it failed to reach the top attack level of “arbitrary code execution” in any test. In simulated enterprise intranet attack tests, Kimi K3 averaged 17 steps completed out of 32, successfully completing the full workflow in only 1 out of 10 attempts.

The report believes that China’s open-weight models continue to make progress on cybersecurity tasks, but there remains a clear gap versus US systems; at the same time, the related results also align with external claims that Kimi K3 may rely on distillation training, leading to insufficient advanced attack capabilities.

View Original
This page may contain third-party content, which is provided for information purposes only (not representations/warranties) and should not be considered as an endorsement of its views by Gate, nor as financial or professional advice. See Disclaimer for details.
  • Reward
  • Comment
  • Repost
  • Share
Comment
Add a comment
Add a comment
No comments
  • Pinned