aiminute. ← All AI news
Research AI Minute Newsroom 2026-08-07

'Approve'? 40,000 game runs show humans wave through one in three dangerous AI agent commands

'Approve'? 40,000 game runs show humans wave through one in three dangerous AI agent commands

Developer-tools firm Scale X turned AI-agent permission prompts into a browser game — and analyzed 40,000+ runs with 409,000 approve/deny decisions. Average accuracy was 66.3%: players approved roughly one in three malicious commands, and a third of sessions ended with a negative score. Vigilance decayed as approvals piled up, and the most-missed threat — an 'npm run analyze' script that can execute arbitrary code — was approved nearly 65% of the time.

Why it matters'A human approves every command' is the industry's standard answer to AI-agent safety. This is the largest public dataset showing that safeguard leaks badly under fatigue — exactly the failure mode regulators and labs assume away. Real protection likely needs sandboxing and better permission design, not more clicking.
#Coding#AI Agents

✓ Verified · 2 sources

WhatsApp X Telegram
Read in the app — free, in 9 languages

Related stories

A model learned to write without the method that trains every AI.
2026-10-06
TikTok put a shopping chatbot inside the video you are watching.
2026-10-06
Reflection will hand out a 501-billion-parameter model for free.
2026-10-06
Wikipedia's owner says OpenAI agents may have caused a May outage.
2026-10-06
Meta's assistant keeps an hourly file on everyone in your life.
2026-10-05