aiminute. ← All AI news
Research 2026-08-17

Two of the world's most decorated mathematicians looked at what AI actually solves, and landed on the same missing piece

Two of the world's most decorated mathematicians looked at what AI actually solves, and landed on the same missing piece

Timothy Gowers, a Fields medallist, published a post on 12 August trying to sort mathematics into the kinds that language models handle well and the kinds they do not. His finding is that models are strong where a result can be reached by assembling known techniques — counterexample hunting is unusually well represented on their hit list — and weak at the judgement of which direction is worth trying at all when the search space is enormous. He is careful to say there is no clean classification yet and asks theorists to help formalise one. Peter Sarnak, writing in the Notices of the American Mathematical Society in July, reached a similar place from a different angle: AI can derive consequences from theory that already exists, but struggles to invent the new abstractions that major proofs are usually built on. A DeepMind paper by Tom Zahavy, 'LLMs Can't Jump', names the same bottleneck in technical terms, calling it the inability to posit new foundational assumptions. None of these are controlled experiments; they are expert assessments of how the tools behave in real mathematical work.

Why it mattersThis is a more useful signal than another benchmark score, because benchmarks mostly measure the thing these mathematicians say models are already good at: closing a gap when the target is stated and the tools are known. The part they identify as missing — deciding what is worth attempting, and inventing the concept that makes a problem tractable — is exactly what does not appear as a question with an answer key, and therefore does not appear in evaluations. It is also a caution against reading Olympiad and competition results as evidence of research-level ability, since competition problems are by construction solvable with known methods.
#Science & Research

✓ Verified · 3 sources

WhatsApp X Telegram
Read in the app — free, in 9 languages

Related stories

Machine learning read the shape of sick brain cells and picked out nine already-approved drugs that calmed them down
2026-08-21
Given four hours and a GPU to improve the way AI is trained, the best agent scored 0.25 out of 1 — and most never tried
2026-08-21
The agent invented a second person to vouch for its code. A 24-year-old in Texas refused to believe either of them.
2026-08-21
Pew put half a million web pages through a detector: a third of everything published since ChatGPT carries its marks
2026-08-21
The FDA has cleared 1,357 AI medical devices. Three of them have been tested on whether patients live longer or better.
2026-08-20