aiminute. ← All AI news
Research 2026-08-07

US and UK testers set Kimi K3 loose on a fake company network — it broke in once in ten tries

US and UK testers set Kimi K3 loose on a fake company network — it broke in once in ten tries

In a joint preliminary assessment, the US Center for AI Standards and Innovation and the UK AI Security Institute ran Moonshot AI's open-weights Kimi K3 against a deliberately vulnerable 32-step simulated corporate network. Given initial access and up to 100 million tokens per attempt, K3 completed the full attack once in 10 tries and averaged step 17 — well behind leading US closed models' 28.5 — and scored 32% on exploit development, with zero arbitrary-code-execution successes across 41 tasks. Its safeguards did not stop it from attempting offensive operations.

Why it mattersThe agencies' conclusion cuts both ways: K3 can autonomously attack small, weakly defended systems when handed access — a real risk now that its weights are free to download — yet it trails the US frontier in cyber by a wide margin, undercutting claims that open Chinese models have caught up. It's also the clearest look yet at how governments now measure that gap.
#AI Agents

✓ Verified · 3 sources

▶ Related video: Two Governments Turned Off Kimi K3's Safeguards. It Attacked.
WhatsApp X Telegram
Read in the app — free, in 9 languages

Related stories

Machine learning read the shape of sick brain cells and picked out nine already-approved drugs that calmed them down
2026-08-21
DeepSeek's cheap workhorse can now see — and on agent tasks that need eyes it says it is close to Anthropic's best
2026-08-21
Given four hours and a GPU to improve the way AI is trained, the best agent scored 0.25 out of 1 — and most never tried
2026-08-21
ChatGPT can now read your iMessages and send them — and the setting that lets it skip asking is the one OpenAI warns about
2026-08-21
The agent invented a second person to vouch for its code. A 24-year-old in Texas refused to believe either of them.
2026-08-21