📄 Research Article
When Help is Unhelpful: Evaluating AI Tutors for Productive Struggle
Synthesis: Zhang et al. (2026) introduce TutorMoments, a replay-based evaluation framework that tests whether LM tutors adapt their pedagogical actions to context — scaffolding when support is needed, pushing for rigor when students are ready, and avoiding over-scaffolding. Evaluating 462 teacher-annotated transcripts from grades 2-7 math tutoring, they find frontier models default toward over-helpfulness at the expense of productive struggle. The paper argues that AI optimized for helpfulness may be misaligned with the pedagogical goal of providing the right help at the right moment.
Summary
TutorMoments evaluates whether LM tutors select instructional actions appropriate to the pedagogical demands of specific learning moments. Expert math teachers annotate key decision points in authentic tutoring transcripts: scaffolding-appropriate moments (student needs support) and rigor-appropriate moments (student is ready for challenge). The framework then replays these moments to test whether LMs select appropriate tutor moves. Findings show minimally prompted LMs frequently over-scaffold and rarely push for rigor; evaluation-aware prompting improves adaptability but behavior concentrates into a narrow set of tutor moves. The dataset includes 462 transcripts from 198 students in grades 2-7 interacting with 173 human tutors, with 1,500+ teacher-annotated key moments.
Key Findings
Connected Concepts
Connected Articles
Citation
Zhang, A., Ross, A., Patel, K., Bernado, J., Bowie, R., Ribeiro, A. T., Halper, D., Valayaputtur, H., Andreas, J., Loeb, S., Lucy, L., Lo, K., & Knight, R. (2026). When Help is Unhelpful: Evaluating AI Tutors for Productive Struggle.