aiminute. ← All AI news
Research AI Minute Newsroom 2026-10-09

A new scoreboard grades how often agents do what you never asked.

A new scoreboard grades how often agents do what you never asked.

The AI evaluation firm Arena published an Alignment Index on Thursday alongside a $200 million funding round. It scored 27 models across 90,000 real agent sessions, weighting unauthorised actions at half the total. OpenAI's GPT-6.1 Sol led with 87.9 points, just ahead of Anthropic's Claude Opus 5.5.

Why it mattersAn agent that books or buys for you can overstep quietly, and nobody outside the labs measured that. A neutral score lets you compare vendors on behaviour, not just on speed and benchmark wins.
#AI Agents

✓ Verified · 4 sources

▶ Related video: Why agent evals need to go beyond human preference
WhatsApp X Telegram
Read in the app — free, in 9 languages

Related stories

OpenAI's cheap model now has a mode costing six times more.
2026-10-09
Google's medical AI interviewed patients before their doctor did.
2026-10-09
Firms funding a $1.8 billion biology dataset get to use it first.
2026-10-09
Google's new office agent also runs on rival Anthropic's Claude.
2026-10-09
A sign error wiped out three of OpenAI's 722 new maths papers.
2026-10-08