Rethinking LLM-Judged Helpfulness as a Pedagogy Signal: A Pre-Registered Audit Across Tutor Models

Created: 2026-07-31 | Tags: llmintelligent-tutoringautomated-gradingbenchmarkfeedback-loop

Shuyi Fan, Boyuan Deng, Mengyu Xu, Jiale Liu, Hongyang Zhang (2026) โ€” arXiv:2607.28128 (cs.CL, cs.AI, cs.CY)

๐Ÿ“„ Full text (arXiv)

Summary

Pre-registered study auditing whether general-purpose helpfulness rubrics can distinguish direct answer-giving from pedagogical guidance in LLM tutors. Uses deterministic detectors for answer leakage and next-turn independent work across three tutor models. Finds that helpfulness ratings conflate genuine pedagogical scaffolding with simply giving correct answers.

The work connects to broader discussions in AI and education around intelligent-tutoring-systems, automated-feedback, llm-evaluation, contributing to our understanding of how llm shapes educational practice.

Key Contributions

Related Pages

Citation

APA: Shuyi Fan, Boyuan Deng, Mengyu Xu, Jiale Liu, Hongyang Zhang (2026). Rethinking LLM-Judged Helpfulness as a Pedagogy Signal: A Pre-Registered Audit Across Tutor Models. arXiv:2607.28128. cs.CL, cs.AI, cs.CY.