📄 Research Article
Confirming Correct, Missing the Rest: LLM Tutoring Agents Struggle Where Feedback Matters Most
Key Finding
LLM tutors achieve near-ceiling on correct steps but systematically over-reject valid-suboptimal reasoning and over-validate incorrect solutions — precisely where adaptive tutoring matters most.
Synthesis
This paper exposes a critical diagnostic blind spot in LLM-based tutoring agents. Across seven models and 10,836 solution-feedback pairs in propositional logic, LLMs performed near-perfectly on clearly correct steps but systematically misfired on the cases that matter most for adaptive tutoring: they over-rejected valid-but-suboptimal reasoning and over-validated incorrect solutions. These failures persisted regardless of solution context, suggesting architectural limitations rather than insufficient information. Alarmingly, even when models correctly diagnosed a step, they often failed to produce pedagogically actionable feedback — revealing a gap between diagnostic accuracy and instructional effectiveness. The authors propose hybrid architectures where knowledge-graph-grounded models handle precise diagnosis while LLMs support open-ended Scaffolding and dialogue. This finding directly complements the behavioral evaluation framework from AI Tutor Behavioral Evaluation, which also found that pedagogical quality alone is insufficient — students must actually act on feedback. Together, these papers suggest that current LLM tutors need both better diagnostic precision AND better feedback-actionability to serve as effective Intelligent Tutoring.
Connected Concepts
Connected Articles
Citation
preprint, A. (2026). Confirming Correct, Missing the Rest: LLM Tutoring Agents Struggle Where Feedback Matters Most