Junyi Yao, Zihao Zheng, Baichuan Li (2026) โ arXiv preprint ๐ Full text (arXiv)
Studies whether public LLM tutoring benchmarks distinguish learning-supportive behavior from mere answer production. Proposes a lightweight diagnostic based on the gap between solving-oriented and pedagogy-oriented benchmark performance. Using MathTutorBench, shows correlation between solving and pedagogy composites is only r=0.421 across 8 models, with several models shifting rank when evaluated on pedagogy. Benchmarks reward guiding questions, calibrated hints, and non-disclosive scaffolding. Recommends reporting solving and pedagogy scores separately.
Key Contributions
- LLM tutoring benchmarks conflate solving with teaching (r=0.421); recommends dual reporting of pedagogy and solving scores.
Related Pages
- agentic-ai-pedagogical-best-practice-2026 โ Automation-learning tension across benchmarks
- ai-k12-evidence-base โ Empirical evidence on AI in K-12 education outcomes
- intelligent-tutoring โ Automated tutoring systems and their evaluation
- student-experience โ Student perspectives on AI in education
- learning-analytics โ Data-driven approaches to understanding learning
- llm-feedback-programming-classroom โ LLM feedback in classroom settings
Citation
APA: Junyi Yao, Zihao Zheng, Baichuan Li (2026). Measuring Whether LLM Tutors Teach or Solve: A Diagnostic for Educational Impact. arXiv:2606.16206. arXiv preprint.