The evaluation firm Robocurve gave three AI models control of a robot arm and five dangerous instructions. OpenAI's GPT-6 Astra refused 2 of 100 trials and completed the dangerous action 60 times. Anthropic's model refused 20 times, and an open research model never refused at all.