aiminute. ← All AI news
Industry AI Minute Newsroom 2026-08-25

Nvidia says its new rack does 30 times more agent work per megawatt — it ran the test itself, and the benchmark's authors have not signed off

Nvidia says its new rack does 30 times more agent work per megawatt — it ran the test itself, and the benchmark's authors have not signed off

On 24 August Nvidia published on-silicon numbers for Vera Rubin NVL72, claiming up to 30x higher throughput per megawatt and up to 35x lower cost per token than the GB300 NVL72 it succeeds. The workload was SemiAnalysis's AgentX — recordings of real agentic coding sessions — run on Kimi K3, MiniMax M3, GLM 5.3, Qwen 3.5 and DeepSeek V4 Pro at contexts above 140,000 tokens. Nvidia says the platform is in full production, and notes in the same post that the results are pending SemiAnalysis review and do not yet reflect Vera CPU performance on tool calls. At Hot Chips the same week it detailed the rack itself: 2 zettaFLOPS of NVFP4 inference and 11 petabytes of HBM4 in a 100MW building, liquid cooling at 45°C, and compute trays with no fans and no cables.

Why it mattersAI capacity is now rationed by electricity rather than by chips — a grid connection is what a data centre waits years for, while hardware ships in months. A 30x gain in work per megawatt, if outside reviewers confirm it, changes how much AI a fixed power contract buys, and that is the number deciding where the next buildings go. Two things are worth keeping in view: the measurement is Nvidia's own, on Nvidia's hardware, before the benchmark's authors reviewed it — and every model in the test is Chinese and open-weight, which is a quiet statement about what the industry now treats as a serious inference workload.
#AI Agents

✓ Verified · 2 sources

WhatsApp X Telegram
Read in the app — free, in 9 languages

Related stories

A model launched last month just brought its maker $870 million.
2026-10-09
One prompt could take over every AI agent in an Amazon account.
2026-10-09
The maker of the agent used on Korean banks deleted its code.
2026-10-09
A new scoreboard grades how often agents do what you never asked.
2026-10-09
Apple cut iPhone 18 Pro orders because AI made memory expensive.
2026-10-09