aiminute. ← All AI news
New Models 2026-08-11

Nvidia's new open model has 30 billion parameters and uses 3 of them at a time — and it runs on a desktop

Nvidia's new open model has 30 billion parameters and uses 3 of them at a time — and it runs on a desktop

Nvidia released Nemotron 3.5 Lightning on 11 August: a 30-billion-parameter mixture-of-experts model that activates only 3 billion parameters per token, with a context window of up to 1 million tokens and weights under the permissive OpenMDW-1.1 licence — companies can download, modify and build on it without asking Nvidia. It is built for the unglamorous half of agent work: code review, tool calls, watching security alerts, answering billing questions, the steps a frontier model would plan but need not perform itself. Nvidia claims up to four times the output speed and 30% faster agentic task completion than models of similar size, and reports 81.62 on MMLU Pro, 75.57 on GPQA Diamond and 52.80 on SWE-bench Verified. It runs locally on RTX PCs, DGX Spark, DGX Station and Jetson boards as well as in data centres, and is available on Hugging Face, ModelScope and OpenRouter. Alongside it Nvidia open-sourced NeMo Switchyard, a routing library that sends each step of a workflow to the cheapest model that can handle it; Nvidia's own benchmarks put a Switchyard-routed setup at close to a third the cost of running Opus 4.8 for everything.

Why it mattersThe router may matter more than the model. Nvidia is arguing that the expensive frontier system should do the planning and almost nothing else — that most of what an agent actually does is small, repetitive work that can be handed to something which fits on a workstation. That is a direct attack on the per-token business of the labs Nvidia sells chips to, and Nvidia shipped both halves of the argument in one day, for free. For anyone running agents on a budget — a school lab, a small team, a solo developer — the practical reading is that the local-model option just gained a serious entry, under a licence that does not require anyone's permission.
#Coding#AI Agents

✓ Verified · 3 sources

WhatsApp X Telegram
Read in the app — free, in 9 languages

Related stories

Rumour: the anonymous model that just topped a coding benchmark, for free, is said to be Zhipu's unreleased flagship
2026-08-21
DeepSeek's cheap workhorse can now see — and on agent tasks that need eyes it says it is close to Anthropic's best
2026-08-21
Nvidia is paying $6 billion for the machine that builds a rival's models — and hiring 109 of the people who ran it
2026-08-21
Given four hours and a GPU to improve the way AI is trained, the best agent scored 0.25 out of 1 — and most never tried
2026-08-21
ChatGPT can now read your iMessages and send them — and the setting that lets it skip asking is the one OpenAI warns about
2026-08-21