Anthropic published a review of its alignment and security work on 31 August. It flagged more than 10 percent of its reinforcement learning environments as broken or open to reward hacking. The company froze changes to those environments for a month and moved 150 engineers to security.