aiminute. ← All AI news
New Models 2026-08-20

The model invents its own tasks, builds the rig to test them, then trains on the results — and DeepReinforce gave the weights away

The model invents its own tasks, builds the rig to test them, then trains on the results — and DeepReinforce gave the weights away

DeepReinforce released Ornith-1.5 on 19 August in three open-weight sizes: a 397-billion-parameter mixture of experts, a 35-billion one that activates 3 billion parameters per token, and a 9-billion dense model, all under an MIT licence with the weights on Hugging Face. The sizes are not the interesting part. Instead of learning from a fixed set of human-written tasks, the model proposes its own, scores each one for validity, difficulty and novelty, writes the scaffolding needed to run it, then generates attempts that feed back into reinforcement learning. The team calls this closing the loop from self-scaffolding to self-improvement. On the lab's own figures the 397B model scores 86.1 on Terminal-Bench 2.1 and 56.0 on DeepSWE, against 85.0 and 59.0 for Claude Opus 4.8 and 82.7 and 46.2 for GLM-5.2, plus 92.8 on GPQA Diamond. Its context window is 262,144 tokens and the full model is roughly 800 gigabytes in BF16, so open here does not mean it runs on your laptop. The 9B version, at 47.0 on Terminal-Bench 2.1 and 70.6 on SWE-bench Verified, is the one that does.

Why it mattersEvery lab now says the supply of good training problems is running short. This is one answer: let the model write the curriculum. If it holds up, the ceiling on capability stops being how many worthwhile problems humans can author and becomes how well a model judges whether a problem it invented is worth solving — a less comfortable bottleneck, because a system that grades its own difficulty can wander into problems that are hard and pointless. These are also DeepReinforce's own numbers on benchmarks it chose, and no independent evaluation has landed yet. What is not in dispute is the licence. An MIT-licensed model claiming parity with a flagship on terminal-agent work is a floor raised in public, and everyone selling access to a closed model now has to price against it.
#Coding#AI Agents

✓ Verified · 4 sources

▶ Related video: Ornith-1.5 Ships 9B Dense, 35B and 397B MoE Models | Models & Agents #Shorts
WhatsApp X Telegram
Read in the app — free, in 9 languages

Related stories

Rumour: the anonymous model that just topped a coding benchmark, for free, is said to be Zhipu's unreleased flagship
2026-08-21
DeepSeek's cheap workhorse can now see — and on agent tasks that need eyes it says it is close to Anthropic's best
2026-08-21
Nvidia is paying $6 billion for the machine that builds a rival's models — and hiring 109 of the people who ran it
2026-08-21
Given four hours and a GPU to improve the way AI is trained, the best agent scored 0.25 out of 1 — and most never tried
2026-08-21
ChatGPT can now read your iMessages and send them — and the setting that lets it skip asking is the one OpenAI warns about
2026-08-21