aiminute. ← All AI news
New Models AI Minute Newsroom 2026-08-27

IBM's new open models were not taught to describe using a terminal. They were trained by using one.

IBM's new open models were not taught to describe using a terminal. They were trained by using one.

IBM released Granite 4.2 on 25 August: three dense reasoning models at 3, 8 and 30 billion parameters, weights published under Apache 2.0, which means download, fine-tune and ship with no licence to negotiate. Each has a thinking mode that can be switched on or off, so the same model can answer a lookup instantly or work a problem step by step. The training is where IBM has put its bet. Beyond ordinary reinforcement learning, the 8B and 30B versions went through what IBM calls agentic RL: rather than learning from written examples of tool use, they were dropped into live sandboxes — real software engineering tasks, a real terminal, real search-driven workflows — and rewarded on whether the job came out right. Coding ability was pushed with a trillion tokens of synthetic code generated by IBM's own CodeAlchemy pipeline. The models take tool calls in the OpenAI format and run on the usual serving stacks; weights are on Hugging Face, Ollama and GitHub, with hosted options on watsonx, OpenRouter, Replicate and others.

Why it mattersThe interesting claim here is not a benchmark, it is a method. Most models learn to use tools the way you might learn to drive from a manual: by reading transcripts of tool use written by someone else. Training inside a live terminal, where the reward is whether the command actually worked, is a different and more expensive proposition — and it is the direction the whole field is drifting, because agentic ability has turned out to depend far more on post-training than on parameter count. For anyone weighing what to run in-house, the practical point is size: a 30B model under Apache 2.0 that has been drilled on real engineering tasks fits on hardware an ordinary company already owns, and comes with no clause about what you may build with it. Treat IBM's own benchmark charts as a starting hypothesis until someone outside IBM reproduces them.
#Coding#AI Agents

✓ Verified · 4 sources

WhatsApp X Telegram
Read in the app — free, in 9 languages

Related stories

A German broker with €60 billion now lets ChatGPT and Claude place your trades. You still have to press yes.
2026-08-27
OpenAI watched its models climb out of the sandbox in May and let the test keep running. In July they had root on a Hugging Face production server.
2026-08-27
Claude's memory now follows you from the chat window into the work app — and it is built to forget your politics unless you say otherwise
2026-08-26
The nameless model that burned through 42 trillion tokens in six days has an owner: Z.ai says Ox Alpha is a GLM, and the weights are opening
2026-08-26
The Qwen4 preview is open: six billion parameters active per token, and on Alibaba's own table it fixes more real bugs than Opus 4.6
2026-08-26