Research
AI Minute Newsroom
2026-08-29
Researchers handed over 70 pieces of their own real lab work as a test. The best AI agent finished 30 percent of it.
Terminal-Bench-Science 0.1, released this week by researchers at Stanford together with the Terminal-Bench and Harbor teams, is a benchmark built from work that scientists actually do. Each of its 70 tasks was contributed by a researcher who took a computational workflow from their own lab and turned it into something an agent can attempt in a terminal — across life, physical, Earth, mathematical and engineering sciences. Every task runs in a container and is checked programmatically, so the score is whether the work came out right, not whether the agent said it did. The results are low by design: Claude Opus 5 leads with a 30.0 percent resolution rate, GPT-5.6 Sol gets 22.4 percent and Claude Fable 5 gets 21.4 percent. Every model scored more than ten points lower than on the general Terminal-Bench 3.0. The project lists 376 contributors across 22 countries and is collecting tasks for version 0.2, with a pull-request deadline of 5 October.
Why it mattersThe claim that AI is about to accelerate science is usually supported by exam-style benchmarks — questions with a known answer, of the kind a scientist stopped being tested on years ago. This measures the other thing: can the system run the pipeline, handle the file formats, notice when a step failed, and produce a result a researcher would accept. Thirty percent is a real number to improve on rather than a saturated one, and because the tasks come from named contributors in specific fields, the failures point somewhere — you can see which sciences the agents are worst at instead of only knowing that they are.
✓ Verified · 4 sources
Read in the app — free, in 9 languages
Related stories
Slack's default model is now Claude, and Salesforce's sales work has been packaged as 37 skills you run from inside the chatbot
2026-08-29Anthropic set Claude loose on ten of its own alignment failures. On the deception task it scored 85 where 28 human safety researchers scored 20.
2026-08-29Z.ai opened GLM-5.3's weights on the day it promised — and quietly dropped the MIT licence its last three models shipped under
2026-08-29The ransomware crew's trick was one sentence: tell the coding agent it is a test. It then spent six weeks inside real company networks.
2026-08-28Google Search will now watch flight prices for you and book the hotel — the conversation ends at a booking, not a list of links
2026-08-28