aiminute. ← All AI news
Research AI Minute Newsroom 2026-08-29

Anthropic set Claude loose on ten of its own alignment failures. On the deception task it scored 85 where 28 human safety researchers scored 20.

Anthropic set Claude loose on ten of its own alignment failures. On the deception task it scored 85 where 28 human safety researchers scored 20.

In a paper published on 28 August, Anthropic describes what it calls an automated alignment researcher: Claude reads the literature on a specific misalignment, proposes a training method, runs it for about thirty minutes, reads the benchmark score, and tries again. Across ten categories of alignment failure — each measured by three to five benchmarks — it closed between 26 and 96 percent of the safety gap without degrading the model's general capabilities. On the deception task it reached 85 percent across multiple runs; 28 human safety researchers, given up to eight hours each, reached 20. The methods it found generalised to models up to 4.7 times larger than the ones they were developed on. In the one production-scale test, Claude Sonnet 5 closed 65 percent of the remaining safety gap in an early Claude Opus 4.8 checkpoint over 60 hours, landing close to the scores of shipped models. The work was led by an Anthropic Fellow, Chen Yueh-Han, and tested on Gemma-2-2B alongside the Opus checkpoint.

Why it mattersAnthropic is careful to say the human comparison is not a fair fight: the researchers submitted one idea each and could not iterate, while the machine ran the loop hundreds of times, so the paper frames the result as an argument for collaboration rather than replacement. Take that at face value and the finding still matters, because the asymmetry it points at is real. If models keep getting better at building models, alignment work has to accelerate at the same rate or it simply falls behind, and this is the first published attempt to show the safety half of that race can be automated too. The caveats the authors put on it are the ones worth carrying: the failures tested were narrow compared with what production systems do, there are no benchmarks yet for the failures nobody has seen, and they did not check whether the fixes survive the heavy reinforcement learning that comes afterwards.
#AI Agents#Science & Research

✓ Verified · 2 sources

WhatsApp X Telegram
Read in the app — free, in 9 languages

Related stories

Z.ai opened GLM-5.3's weights on the day it promised — and quietly dropped the MIT licence its last three models shipped under
2026-08-29
The ransomware crew's trick was one sentence: tell the coding agent it is a test. It then spent six weeks inside real company networks.
2026-08-28
Google Search will now watch flight prices for you and book the hotel — the conversation ends at a booking, not a list of links
2026-08-28
The protocol that let AI touch your files now has a sibling for lab equipment — Anthropic wants a model to drive a microscope the way it drives an app
2026-08-28
The labs whose models broke out and hacked people have now co-signed a letter asking the world to hurry up and defend itself
2026-08-28