The Missing Evaluation Axis: What 10,000 Student Submissions Reveal About AI Tutor Effectiveness

Created: 2026-05-09 | Tags: intelligent-tutoringefficacy-studyhigher-edbenchmarkengagement-metrics
๐Ÿ“„ Full text: arXiv:2605.05648 ยท local

Definition

A framework for evaluating AI tutoring systems that extends beyond pedagogical quality of feedback to measure what students actually do with that feedback โ€” whether they act on it and whether they apply it correctly. Proposed by Niousha et al. (2026) based on analysis of 10,235 real student code submissions.

Key Findings

Significance for AI in Education

This work addresses a critical evaluation gap. An AI tutor that gives perfect pedagogical feedback is worthless if students ignore it or apply it incorrectly. The behavioral axis complements pedagogical assessment to provide a complete picture of real-world effectiveness. This has direct implications for ai-tutor-effectiveness-review and challenges the assumptions in tutoring-specific-vs-general-ai about what makes tutoring effective.

Connections

Open Questions

Related Pages