Catching The Correct Answer Trap: Characterising AI Tutor Blind Spots When Analysing Student Reasoning

Created: 2026-05-28 | Tags: intelligent-tutoringautomated-gradingformative-assessmentllmscaffolding

Imran et al. (2026) โ€” University College London / Eedi. AIED 2026.

๐Ÿ“„ Full text (arXiv)

Catching the Correct Answer Trap โ€” accepted at AIED 2026 โ€” exposes a critical blind spot in intelligent-tutoring systems: they systematically fail to detect misconceptions when students arrive at correct answers through flawed reasoning. Using real student data from the Eedi mathematics platform, the authors characterize the 'Correct Answer Trap' (CAT), showing that 71% of failures concentrate in just two question types where erroneous reasoning accidentally produces the correct numerical answer. Even a frontier llm achieves only 84% detection accuracy while generating roughly 4 false alarms per genuine detection โ€” making standalone automated screening impractical. This finding has profound implications for automated-grading and formative-assessment systems: high overall accuracy metrics can mask catastrophic failures in reasoning assessment. The work connects to llm-student-misconception-identification research on the gap between answer checking and reasoning evaluation, and to codify-socratic-programming-tutor findings that even Socratic AI tutors can miss deep misconceptions. The paper reinforces calls for human-in-the-loop approaches in intelligent-tutoring and suggests that scaffolding designs should explicitly account for reasoning assessment, not just answer verification. The concentration of failures in predictable question types also suggests targeted improvements are possible.

Related Pages

Citation

APA: Moiz Imran, Sahan Bulathwela (2026). Catching The Correct Answer Trap: Characterising AI Tutor Blind Spots When Analysing Student Reasoning. arXiv:2605.23925. AIED 2026.