📄 Research Article
Measuring Whether LLM Tutors Teach or Solve: A Diagnostic for Educational Impact
Studies whether public LLM tutoring benchmarks distinguish learning-supportive behavior from mere answer production. Proposes a lightweight diagnostic based on the gap between solving-oriented and pedagogy-oriented benchmark performance. Using MathTutorBench, shows correlation between solving and pedagogy composites is only r=0.421 across 8 models, with several models shifting rank when evaluated on pedagogy. Benchmarks reward guiding questions, calibrated hints, and non-disclosive scaffolding. Recommends reporting solving and pedagogy scores separately.
Key Findings
Study Design & Method
The diagnostic exploits the fact that public tutoring benchmarks (MathTutorBench, TutorBench) score models on multiple rubrics. By separating rubric items into solving-oriented and pedagogy-oriented composites, the authors compute a per-model gap that reveals whether a model's benchmark standing reflects teaching quality or merely answer production. The correlational analysis across eight models quantifies how partially aligned the two dimensions are, while the rubric analysis identifies which specific behaviors — guiding questions, calibrated hints, non-disclosive scaffolding — benchmarks already reward.
Implications for AI in Education
For the Benchmark community and for AI tutor deployment, the findings argue for reporting solving-oriented and pedagogy-oriented scores separately and for making disclosure-sensitive, student-agency-preserving criteria more explicit. A model that tops a solving leaderboard should not be assumed to be a good tutor; evaluation infrastructure must measure learning support directly. This connects to Scaffolding and to the design of AI Tutoring systems where the goal is not the fastest answer but durable student understanding.
Connected Concepts
Connected Articles
Citation
Junyi Yao, Zihao Zheng, Baichuan Li (2026). Measuring Whether LLM Tutors Teach or Solve: A Diagnostic for Educational Impact. arXiv:2606.16206. arXiv preprint.