aiminute. ← All AI news
Research AI Minute Newsroom 2026-09-13

The best coding agent finished under four in ten real company tasks.

The best coding agent finished under four in ten real company tasks.

Specific Labs published Real-SWE, a test built from private production code licensed from real companies. Eight frontier models solved between 16.2 and 38.8 percent of the tasks, with Claude Fable 5.1 on top. GPT-6 Astra reached 33.8 percent and Gemini 3.8 Flash 31.2 percent on the same set.

Why it mattersBenchmark scores you read about come from tidy public repositories, not messy company code. This is the gap between a demo and the codebase your team actually maintains.
#Coding#AI Agents

✓ Verified · 2 sources

WhatsApp X Telegram
Read in the app — free, in 9 languages

Related stories

Anthropic's boss says AI swarms could own the internet within a year.
2026-09-13
A top-severity flaw let anyone run code in Google's agent kit.
2026-09-12
Salesforce gave its agents human names and weeks to finish a job.
2026-09-12
25 Fields medallists signed one letter against AI problem-solving.
2026-09-12
Kimi's cheap coding model now opens a million-token window to all.
2026-09-12