aiminute. ← All AI news
Research AI Minute Newsroom 2026-10-06

Twenty tries at the same task cut the best agent's score by a third.

Twenty tries at the same task cut the best agent's score by a third.

A new benchmark called Thinkingbox ran 507 business workflows twenty times each from a clean start. Claude Opus 5 solved 66.5 percent of tasks once, and 47.5 percent on all twenty runs. Kimi K3 fell harder, from 57.4 percent down to 17.6 percent, so reliability is the gap.

Why it mattersDemos show an agent succeeding once; your work needs it to succeed every time. This test grades the database the agent left behind, not the summary it wrote.
#AI Agents

✓ Verified · 1 sources

WhatsApp X Telegram
Read in the app — free, in 9 languages

Related stories

A model learned to write without the method that trains every AI.
2026-10-06
TikTok put a shopping chatbot inside the video you are watching.
2026-10-06
Wikipedia's owner says OpenAI agents may have caused a May outage.
2026-10-06
Meta's assistant keeps an hourly file on everyone in your life.
2026-10-05
Two senators want prison time for bosses whose AI agents hack.
2026-10-05