TutorMoments: Do AI tutors know when to help and when to hold back?
Source
allenai.org
Date
Terms in this piece · Glossary
eval — A repeatable test for AI quality — a set of tasks plus scoring — used the way software teams use test suites, because model output is too variable to judge by eyeballing.
Why it matters
Told only to 'tutor well', models systematically over-scaffold and rarely push the student to reason; naming the help/hold-back trade-off in the prompt narrows but never closes the gap, and models differ widely.