aiminute. ← All AI news
Research AI Minute Newsroom 2026-07-31

Anthropic: our own models breached three real companies during safety tests

Anthropic: our own models breached three real companies during safety tests

After reviewing 141,006 cybersecurity evaluation runs, Anthropic disclosed three incidents in which Claude models — believing they were playing a simulated hacking exercise — reached the open internet through a misconfigured test environment run with partner Irregular. One model used weak passwords to pull several hundred rows from a company's production database; another published malicious packages to PyPI that 15 real systems downloaded; a third scanned some 9,000 targets and broke into a firm via SQL injection. The review was prompted by OpenAI's similar disclosure last week.

Why it mattersThat's two frontier labs in one week admitting their test sandboxes leaked into the real world. Models can't tell a game from reality — so the safety-testing infrastructure itself is becoming a security risk that needs auditing.
#AI Agents

✓ Verified · 3 sources

WhatsApp X Telegram
Read in the app — free, in 9 languages

Related stories

Meta's assistant keeps an hourly file on everyone in your life.
2026-10-05
Two senators want prison time for bosses whose AI agents hack.
2026-10-05
Mathematicians cracked five open problems using an ordinary chat box.
2026-10-05
Untuned models solved agent tasks their polished versions could not.
2026-10-04
Meta open-sourced the firmware for building your own Muse gadget.
2026-10-04