Research
2026-08-22
The model on its own scored 30 percent. Wrapped in Nvidia's scaffolding, the same model cleared everything.
Nvidia published results on 21 August for AVO, short for Agentic Variation Operators, a general-purpose coding agent system that completed all 183 levels across all 25 environments in the public set of ARC-AGI-3, the interactive reasoning benchmark, scoring 100.00 on the benchmark's measure of relative human action efficiency. The environments hand an agent no instructions, no stated rules and no goal; it has to work out what the game is by playing it. The language model doing the thinking was Anthropic's Claude Opus 5, which manages roughly 30 percent on the same set at high reasoning effort when run on its own. AVO needed 6,624 environment actions, about 12 percent fewer than VISTA, the previous best. Nvidia notes the run covers only the 25-environment public set, not the semi-private or private competition sets.
Why it mattersFor two years the working assumption has been that progress arrives with the next model. This result says a large share of it is now available from the software wrapped around a model you can already buy: memory, planning, supervision, a loop that inspects its own work and tries again. The same weights went from a third of the benchmark to all of it without being retrained. That moves where the engineering effort belongs, and it makes the question of which model a product uses a much weaker predictor of what that product can actually do.
✓ Verified · 3 sources
▶ Related video: Interactive Reasoning Benchmarks | ARC-AGI-3 Preview
Read in the app — free, in 9 languages
Related stories
Three-quarters of Americans do not want a data centre near them. A year ago they were evenly split.
2026-08-22Anthropic will let its strongest model hunt flaws in your code — but it will not let you talk to it
2026-08-22Machine learning read the shape of sick brain cells and picked out nine already-approved drugs that calmed them down
2026-08-21Rumour: the anonymous model that just topped a coding benchmark, for free, is said to be Zhipu's unreleased flagship
2026-08-21DeepSeek's cheap workhorse can now see — and on agent tasks that need eyes it says it is close to Anthropic's best
2026-08-21