A paper posted on 1 October compared 14 pairs of models, each a raw version and its tuned twin. With enough attempts, raw models cracked tasks the tuned ones never solved at all. The authors call the gap a sharpening tax and offer a sampling fix for it.