aiminute. ← All AI news
Research AI Minute Newsroom 2026-09-02

Anthropic found faults in over 10 percent of its own training tasks.

Anthropic found faults in over 10 percent of its own training tasks.

Anthropic published a review of its alignment and security work on 31 August. It flagged more than 10 percent of its reinforcement learning environments as broken or open to reward hacking. The company froze changes to those environments for a month and moved 150 engineers to security.

Why it mattersModels pick up habits in these practice environments, so a broken one teaches the wrong thing. It is rare for a lab to count its own faults in public.

✓ Verified · 2 sources

WhatsApp X Telegram
Read in the app — free, in 9 languages

Related stories

OpenAI's next model thinks in loops that leave fewer traces.
2026-09-03
Perplexity cites 215,128 buying guides made by three linked websites.
2026-09-02
One developer beat many chatbots on a reasoning test for 67 cents.
2026-09-02
The best AI agent finished 30% of real lab work
2026-08-29
Claude beat human researchers at closing alignment gaps
2026-08-29