---
source_url: https://arxiv.org/abs/2605.05648
ingested: 2026-05-09
sha256: df1297f15c57525bee040e3866d40fa042488b4c2d4e19f3ed4f2d83936eedd1
---

# The Missing Evaluation Axis: What 10,000 Student Submissions Reveal About AI Tutor Effectiveness

**Authors:** Rose Niousha, Samantha Boatright Smith, Bita Akram, Peter Brusilovsky, Arto Hellas, Juho Leinonen, John DeNero, Narges Norouzi
**Published:** 2026-05-07
**Venue:** AIED 2026
**URL:** https://arxiv.org/abs/2605.05648

## Abstract
Current AI tutoring systems are primarily evaluated on pedagogical quality of feedback. This paper argues for adding a behavioral evaluation dimension grounded in student interaction data. Analysis of 10,235 code submissions with AI tutor feedback reveals substantial differences in student engagement patterns invisible to pedagogy-only evaluation. Behavioral signals (whether students act on feedback and apply it correctly) more strongly predict perceived helpfulness than pedagogical quality alone.
