Research
AI Minute Newsroom
2026-08-25
A Stanford lab has begun a three-month, 535-billion-parameter training run in public — loss curves, data mix and failed experiments included
Percy Liang's Marin project has started pretraining Marin 535B-A23B, a mixture-of-experts model with 535 billion parameters and 23 billion active per token, on 11 GB200 NVL72 racks. The published plan is 18.75 trillion tokens — 80 percent pretraining, 20 percent midtraining — over roughly three months and about 2.7e24 FLOPs, with post-training to follow. What separates it from a normal release is that the run itself is open: live loss curves, the data mixture, hardware telemetry, configurations, engineering artifacts and the experiments that did not work, with outsiders able to propose and run experiments through GitHub. Before starting, the team trained a ladder of smaller models to fix the recipe.
Why it mattersAlmost everything the public knows about training a large model comes from finished weights and a paper written afterwards. The judgement calls, the dead ends and the hours when a run destabilises are exactly what labs do not publish, and they are most of the actual craft. A run at 2.7e24 FLOPs happening in the open is the first chance for people outside a frontier lab to watch those decisions as they are taken — at a scale close enough to the frontier that what is learned still transfers.
✓ Verified · 3 sources
Read in the app — free, in 9 languages
Related stories
A new scoreboard grades how often agents do what you never asked.
2026-10-09Google's medical AI interviewed patients before their doctor did.
2026-10-09Firms funding a $1.8 billion biology dataset get to use it first.
2026-10-09A sign error wiped out three of OpenAI's 722 new maths papers.
2026-10-08Buterin says AI maths could break crypto before quantum computers do.
2026-10-08