Research
AI Minute Newsroom
2026-08-17
Two of the world's most decorated mathematicians looked at what AI actually solves, and landed on the same missing piece
Timothy Gowers, a Fields medallist, published a post on 12 August trying to sort mathematics into the kinds that language models handle well and the kinds they do not. His finding is that models are strong where a result can be reached by assembling known techniques — counterexample hunting is unusually well represented on their hit list — and weak at the judgement of which direction is worth trying at all when the search space is enormous. He is careful to say there is no clean classification yet and asks theorists to help formalise one. Peter Sarnak, writing in the Notices of the American Mathematical Society in July, reached a similar place from a different angle: AI can derive consequences from theory that already exists, but struggles to invent the new abstractions that major proofs are usually built on. A DeepMind paper by Tom Zahavy, 'LLMs Can't Jump', names the same bottleneck in technical terms, calling it the inability to posit new foundational assumptions. None of these are controlled experiments; they are expert assessments of how the tools behave in real mathematical work.
Why it mattersThis is a more useful signal than another benchmark score, because benchmarks mostly measure the thing these mathematicians say models are already good at: closing a gap when the target is stated and the tools are known. The part they identify as missing — deciding what is worth attempting, and inventing the concept that makes a problem tractable — is exactly what does not appear as a question with an answer key, and therefore does not appear in evaluations. It is also a caution against reading Olympiad and competition results as evidence of research-level ability, since competition problems are by construction solvable with known methods.
✓ Verified · 3 sources
Read in the app — free, in 9 languages
Related stories
A model learned to write without the method that trains every AI.
2026-10-06The US got 16 countries to put AI at the centre of science.
2026-10-05Mathematicians cracked five open problems using an ordinary chat box.
2026-10-05Untuned models solved agent tasks their polished versions could not.
2026-10-04Pretraining on everyday photos got a model to 70% on a reasoning test.
2026-10-04