AI in Education
2026-08-09
AI tutors hand over the answer: teachers scored seven models, and almost none of them pushed back
The Allen Institute for AI released TutorMoments on August 7, a benchmark built from 462 de-identified one-on-one math tutoring transcripts with US students in grades 2 to 7, drawn from a high-dosage tutoring programme serving Title I schools. Experienced maths teachers marked more than 1,500 moments in those transcripts where a tutor had to choose: make the problem easier to start, or push the student to do the thinking. Seven models were then replayed into those moments — Gemini 2.5 Pro, Gemini 3.5 Flash, Claude Opus 4.8, Claude Sonnet 4.6, GPT-5.5, GPT-5.4 mini and DeepSeek V4 Pro. Asked plainly, without being told what good tutoring looks like, the models scored as low as 0.023 to 0.046 on rigour: they almost never asked the student to work. Told explicitly what was being measured, every model improved sharply — Gemini 2.5 Pro reached 0.896 on appropriate scaffolding, Claude Opus 4.8 0.831 on rigour, Claude Sonnet 4.6 0.913 on avoiding over-help — but gaps remained. Ai2 published the tech report, the dataset and the code.
Why it mattersThis measures the thing parents and teachers actually worry about, and it puts a number on it. A chatbot's default instinct is to be helpful, and in tutoring, being maximally helpful is the failure mode: the student leaves with a finished problem and no new skill. The near-zero rigour scores on the plain prompt say the default setting of every major model is 'here, let me'. The good news is the second half of the result — the same models get much better when told to hold back. That is a prompt, a system message, a product decision. So if your child or your class is using a general chatbot as a tutor, the instruction to make it teach rather than solve is not a nice extra; it is the whole difference.
✓ Verified · 2 sources
Read in the app — free, in 9 languages
Related stories
For two years schools tried to keep chatbots out. Now they are handing them to pupils and asking them to catch the mistakes.
2026-08-21Two school years, 18 schools, one AI tutor: the gain was about what practising without it produces
2026-08-20A university system pointed a model at 14,000 syllabuses and asked it to find the forbidden topics
2026-08-19ChatGPT now guesses whether it is talking to a child — and if it decides yes, the product changes around them
2026-08-18Mexico's largest university is making tens of thousands sit its entrance exam again — in a room, on paper, with nothing in their pockets
2026-08-17