Developer-tools firm Scale X turned AI-agent permission prompts into a browser game — and analyzed 40,000+ runs with 409,000 approve/deny decisions. Average accuracy was 66.3%: players approved roughly one in three malicious commands, and a third of sessions ended with a negative score. Vigilance decayed as approvals piled up, and the most-missed threat — an 'npm run analyze' script that can execute arbitrary code — was approved nearly 65% of the time.