Today we're introducing a preview of TutorMoments, a framework that measures whether AI tutors can make one of the hardest calls in teaching: when to step in and help a student, & when to hold back and let them do the heavy thinking. 🧵

Language models are trained to be helpful, and a helpful assistant tends to do the hard part of learning for you: explains the concept, lays out the steps, & guides you to the answer. That can cut short the productive struggle that leads to stronger understanding.
TutorMoments is built on transcripts of real one-on-one math tutoring. We had experienced teachers read them & flag key moments—decision points where the tutor had to choose between making a problem easier & pushing the student to do more of the reasoning.
TutorMoments pauses a transcript at these key moments & lets an LLM take over as the tutor; another model stands in for the student. Each replay is scored: did the tutor support the student when needed, push for harder thinking when they were ready, & avoid over-helping?
Default helpfulness training makes models poor tutors: they do the reasoning for the student. TutorMoments scores that tradeoff at teacher-marked decision points and shows prompting alone improves every model but does not close the gap.
Checking sign-in…
Loading comments…