aiminute. ← All AI news
Research AI Minute Newsroom 2026-08-25

A Stanford lab has begun a three-month, 535-billion-parameter training run in public — loss curves, data mix and failed experiments included

A Stanford lab has begun a three-month, 535-billion-parameter training run in public — loss curves, data mix and failed experiments included

Percy Liang's Marin project has started pretraining Marin 535B-A23B, a mixture-of-experts model with 535 billion parameters and 23 billion active per token, on 11 GB200 NVL72 racks. The published plan is 18.75 trillion tokens — 80 percent pretraining, 20 percent midtraining — over roughly three months and about 2.7e24 FLOPs, with post-training to follow. What separates it from a normal release is that the run itself is open: live loss curves, the data mixture, hardware telemetry, configurations, engineering artifacts and the experiments that did not work, with outsiders able to propose and run experiments through GitHub. Before starting, the team trained a ladder of smaller models to fix the recipe.

Why it mattersAlmost everything the public knows about training a large model comes from finished weights and a paper written afterwards. The judgement calls, the dead ends and the hours when a run destabilises are exactly what labs do not publish, and they are most of the actual craft. A run at 2.7e24 FLOPs happening in the open is the first chance for people outside a frontier lab to watch those decisions as they are taken — at a scale close enough to the frontier that what is learned still transfers.

✓ Verified · 3 sources

WhatsApp X Telegram
Read in the app — free, in 9 languages

Related stories

State-linked hacking crews more than doubled their attack volume once they started handing the boring parts to a model
2026-08-25
A question mathematicians have not been able to answer since 1948 was answered on Sunday, in a 108-page file posted to X
2026-08-25
By the end of last year, nine in ten biomedical papers carried the fingerprints of a language model
2026-08-24
Roughly 90 percent of executives say AI has not raised productivity. The cuts are happening anyway.
2026-08-24
An older Claude broke Anthropic's own content rule in 10 tries out of 10 — and it is still sold through Azure and Bedrock
2026-08-23