📄 Research Article
Rethinking LLM-Judged Helpfulness as a Pedagogy Signal: A Pre-Registered Audit Across Tutor Models
Pre-registered study auditing whether general-purpose helpfulness rubrics can distinguish direct answer-giving from pedagogical guidance in LLM tutors. Uses deterministic detectors for answer leakage and next-turn independent work across three tutor models. Finds that helpfulness ratings conflate genuine pedagogical scaffolding with simply giving correct answers.
Key Findings
Study Design & Method
The audit uses deterministic process measures — detectors for answer leakage and next-turn independent work — alongside LLM judges, over 1,179 confirmatory answer-phase tutor turns. Each session follows a fixed five-phase protocol: six training problems, then immediate, interference, and delayed probes, plus transfer probes; within a training problem the tutor and student alternate for a fixed number of turns. Because the two policies share the same underlying model and student, any difference is attributable to the policy itself, making the design a controlled test of whether helpfulness rubrics carry pedagogical signal.
Implications for AI in Education
The central conclusion is that general-purpose helpfulness is not a reliable pedagogy signal in this controlled setting: a rubric tuned to "helpful" answers cannot distinguish a tutor that scaffolds from one that leaks the answer. Tutor evaluation should therefore pair pedagogy-targeted rubrics with deterministic process measures such as answer leakage and next-turn independent work. For AI Tutoring and Benchmark design, this argues against relying on preference-based helpfulness judgments and toward measurement of student agency and independent work.
Connected Concepts
Connected Articles
Citation
Shuyi Fan, Boyuan Deng, Mengyu Xu, Jiale Liu, Hongyang Zhang (2026). Rethinking LLM-Judged Helpfulness as a Pedagogy Signal: A Pre-Registered Audit Across Tutor Models. arXiv:2607.28128. cs.CL, cs.AI, cs.CY.