aiminute. ← All AI news
New Models 2026-08-15

The company behind China's Instagram open-sourced the sibling of the model that got a perfect score at the Maths Olympiad

The company behind China's Instagram open-sourced the sibling of the model that got a perfect score at the Maths Olympiad

Xiaohongshu — the lifestyle app known abroad as RedNote — released open weights on 14 August for dots3-note-prev, a mixture-of-experts model with 280 billion total parameters but only 16 billion active at a time, a 512K context window, and text, vision and speech inputs. It comes from the same family as dots-note-3.0, which last month recorded what its lab describes as an AI's first official perfect 42 out of 42 at the International Mathematical Olympiad. The lab also published its training method, TEMPO, in which the model alternates between acting and grading its own progress on long tasks, plus two benchmarks built for exactly those tasks: VibeSearchBench, 200 tasks across 20 domains, and VibeLifeBench, 20 tasks with 1,247 individual checks. The lab reports that neither Claude Opus 5 nor GPT-5.5 cleared its pass mark on them.

Why it mattersTwo things are unusual here. The first is who: a social shopping app, not a lab, is now shipping open weights in the same week as Alibaba and Meta — a reminder that in China the consumer internet companies have their own frontier teams and are willing to give the results away. The second is what they measured. Almost every public benchmark rewards a model for answering a question; these two reward it for finishing a multi-hour errand — planning a trip, running an operation — and grade it on more than a thousand small checks. That is the harder and more honest test of an agent, and the claim that two frontier models fail it deserves independent replication before anyone treats it as settled.
#AI Agents

✓ Verified · 4 sources

WhatsApp X Telegram
Read in the app — free, in 9 languages

Related stories

DeepSeek's cheap workhorse can now see — and on agent tasks that need eyes it says it is close to Anthropic's best
2026-08-21
Given four hours and a GPU to improve the way AI is trained, the best agent scored 0.25 out of 1 — and most never tried
2026-08-21
ChatGPT can now read your iMessages and send them — and the setting that lets it skip asking is the one OpenAI warns about
2026-08-21
Show the robot once — three to twelve seconds — and it gets the job right 59 times out of 100 with no training at all
2026-08-21
The agent invented a second person to vouch for its code. A 24-year-old in Texas refused to believe either of them.
2026-08-21