aiminute. ← All AI news
Research AI Minute Newsroom 2026-08-07

'Approve'? 40,000 game runs show humans wave through one in three dangerous AI agent commands

'Approve'? 40,000 game runs show humans wave through one in three dangerous AI agent commands

Developer-tools firm Scale X turned AI-agent permission prompts into a browser game — and analyzed 40,000+ runs with 409,000 approve/deny decisions. Average accuracy was 66.3%: players approved roughly one in three malicious commands, and a third of sessions ended with a negative score. Vigilance decayed as approvals piled up, and the most-missed threat — an 'npm run analyze' script that can execute arbitrary code — was approved nearly 65% of the time.

Why it matters'A human approves every command' is the industry's standard answer to AI-agent safety. This is the largest public dataset showing that safeguard leaks badly under fatigue — exactly the failure mode regulators and labs assume away. Real protection likely needs sandboxing and better permission design, not more clicking.
#Coding#AI Agents

✓ Verified · 2 sources

WhatsApp X Telegram
Read in the app — free, in 9 languages

Related stories

Meta's assistant keeps an hourly file on everyone in your life.
2026-10-05
Two senators want prison time for bosses whose AI agents hack.
2026-10-05
OpenAI will ship a Codex upgrade daily for 28 days or reset limits.
2026-10-05
Mathematicians cracked five open problems using an ordinary chat box.
2026-10-05
Untuned models solved agent tasks their polished versions could not.
2026-10-04