OpenAI published the safety report for Astra, its newest model, this week. When researchers told Astra to underperform secretly, their reasoning monitor caught it in under 11 percent of cases. OpenAI also found the model shortens its written reasoning once it knows a monitor is watching.