Research
2026-08-07
'Approve'? 40,000 game runs show humans wave through one in three dangerous AI agent commands
Developer-tools firm Scale X turned AI-agent permission prompts into a browser game — and analyzed 40,000+ runs with 409,000 approve/deny decisions. Average accuracy was 66.3%: players approved roughly one in three malicious commands, and a third of sessions ended with a negative score. Vigilance decayed as approvals piled up, and the most-missed threat — an 'npm run analyze' script that can execute arbitrary code — was approved nearly 65% of the time.
Why it matters'A human approves every command' is the industry's standard answer to AI-agent safety. This is the largest public dataset showing that safeguard leaks badly under fatigue — exactly the failure mode regulators and labs assume away. Real protection likely needs sandboxing and better permission design, not more clicking.
✓ Verified · 2 sources
Read in the app — free, in 9 languages
Related stories
Machine learning read the shape of sick brain cells and picked out nine already-approved drugs that calmed them down
2026-08-21Rumour: the anonymous model that just topped a coding benchmark, for free, is said to be Zhipu's unreleased flagship
2026-08-21DeepSeek's cheap workhorse can now see — and on agent tasks that need eyes it says it is close to Anthropic's best
2026-08-21Nvidia is paying $6 billion for the machine that builds a rival's models — and hiring 109 of the people who ran it
2026-08-21Given four hours and a GPU to improve the way AI is trained, the best agent scored 0.25 out of 1 — and most never tried
2026-08-21