aiminute. ← All AI news
Research 2026-08-16

They built a language model that has never read anything above fifth grade — then found they could not teach it more

They built a language model that has never read anything above fifth grade — then found they could not teach it more

Researchers at the Max Planck Institute for Intelligent Systems, the ELLIS Institute Tübingen and ETH Zürich released LittleLearner: language models of 0.6B, 1.3B and 5B parameters trained from scratch on an 88-billion-token corpus distilled from FineWeb-Edu through a five-stage filter aligned to US Common Core standards, with everything taught above Grade 5 removed. Each model has a matching control trained on the same pipeline without the filter. Their central finding is that scaling the model up, post-training with SFT and GRPO, and in-context learning all amplify what the curriculum already contained — but none of them meaningfully improved performance on knowledge outside it. The corpus and models are published as a sandbox for studying how models acquire knowledge.

Why it mattersAlmost every claim about what a model "learned" is untestable, because nobody can say what was in the training data. This one has a deliberate edge, which turns it into a controlled laboratory for an argument the field keeps having: do the standard techniques for improving models actually add knowledge, or do they just extract more of what was already there? Within this boundary the answer was the second one — and that is a result you cannot get by poking at a frontier model whose diet is unknown.
#Science & Research

✓ Verified · 2 sources

WhatsApp X Telegram
Read in the app — free, in 9 languages

Related stories

Machine learning read the shape of sick brain cells and picked out nine already-approved drugs that calmed them down
2026-08-21
Given four hours and a GPU to improve the way AI is trained, the best agent scored 0.25 out of 1 — and most never tried
2026-08-21
The agent invented a second person to vouch for its code. A 24-year-old in Texas refused to believe either of them.
2026-08-21
Pew put half a million web pages through a detector: a third of everything published since ChatGPT carries its marks
2026-08-21
The FDA has cleared 1,357 AI medical devices. Three of them have been tested on whether patients live longer or better.
2026-08-20