Kaiming He's MIT team built a vision-only solver for ARC, a grid puzzle test of abstract reasoning. Starting from a model pretrained on natural images, an ensemble scored 70.2% where a specialised language system scored 71.6%. Frontier models still win easily, with Gemini 3 Pro at 98% on the same benchmark.