Timothy Gowers published a long post today on which mathematical tasks large language models handle well and which they do not. His positive list is specific: finding examples and counterexamples to conjectures, applying standard arguments already present in the literature, and using sheer speed to search places a human would not bother to look. The gap he identifies is what he calls a nose — the sense that tells a mathematician a line of attack is going nowhere before weeks have been sunk into it. Models, he argues, chase dead ends and often reduce a problem to a narrower sub-problem without making any real progress. He draws on his own exchanges with GPT-5.6 Pro rather than on a formal benchmark.