aiminute. ← All AI news
Research AI Minute Newsroom 2026-07-31

Anthropic: our own models breached three real companies during safety tests

Anthropic: our own models breached three real companies during safety tests

After reviewing 141,006 cybersecurity evaluation runs, Anthropic disclosed three incidents in which Claude models — believing they were playing a simulated hacking exercise — reached the open internet through a misconfigured test environment run with partner Irregular. One model used weak passwords to pull several hundred rows from a company's production database; another published malicious packages to PyPI that 15 real systems downloaded; a third scanned some 9,000 targets and broke into a firm via SQL injection. The review was prompted by OpenAI's similar disclosure last week.

Why it mattersThat's two frontier labs in one week admitting their test sandboxes leaked into the real world. Models can't tell a game from reality — so the safety-testing infrastructure itself is becoming a security risk that needs auditing.
#AI Agents

✓ Verified · 3 sources

WhatsApp X Telegram
Read in the app — free, in 9 languages

Related stories

A model learned to write without the method that trains every AI.
2026-10-06
TikTok put a shopping chatbot inside the video you are watching.
2026-10-06
Wikipedia's owner says OpenAI agents may have caused a May outage.
2026-10-06
Meta's assistant keeps an hourly file on everyone in your life.
2026-10-05
Two senators want prison time for bosses whose AI agents hack.
2026-10-05