📄 Research Article
The Missing Evaluation Axis: What 10,000 Student Submissions Reveal About AI Tutor Effectiveness
A framework for evaluating AI tutoring systems that extends beyond pedagogical quality of feedback to measure what students actually do with that feedback — whether they act on it and whether they apply it correctly. Proposed by Niousha et al. (2026) based on analysis of 10,235 real student code submissions.
AI Tutor Behavioral Evaluation
Definition
A framework for evaluating AI tutoring systems that extends beyond pedagogical quality of feedback to measure what students actually do with that feedback — whether they act on it and whether they apply it correctly. Proposed by Niousha et al. (2026) based on analysis of 10,235 real student code submissions.
Key Findings
Significance for AI in Education
This work addresses a critical evaluation gap. An AI tutor that gives perfect pedagogical feedback is worthless if students ignore it or apply it incorrectly. The behavioral axis complements pedagogical assessment to provide a complete picture of real-world effectiveness. This has direct implications for AI Tutor Effectiveness Review and challenges the assumptions in Tutoring Specific Vs General AI about what makes tutoring effective.
Open Questions
Connected Concepts
Connected Articles
Citation
Niousha, R., Smith, S.B., Akram, B., Brusilovsky, P., Hellas, A., Leinonen, J., DeNero, J., & Norouzi, N. (2026). The Missing Evaluation Axis: What 10,000 Student Submissions Reveal About AI Tutor Effectiveness