Measuring Whether LLM Tutors Teach or Solve: A Diagnostic for Educational Impact

Created: 2026-06-16 | Tags: intelligent-tutoringllmfeedback-loopscaffoldingbenchmark

Junyi Yao, Zihao Zheng, Baichuan Li (2026) โ€” arXiv preprint ๐Ÿ“„ Full text (arXiv)

Studies whether public LLM tutoring benchmarks distinguish learning-supportive behavior from mere answer production. Proposes a lightweight diagnostic based on the gap between solving-oriented and pedagogy-oriented benchmark performance. Using MathTutorBench, shows correlation between solving and pedagogy composites is only r=0.421 across 8 models, with several models shifting rank when evaluated on pedagogy. Benchmarks reward guiding questions, calibrated hints, and non-disclosive scaffolding. Recommends reporting solving and pedagogy scores separately.

Key Contributions

Related Pages

Citation

APA: Junyi Yao, Zihao Zheng, Baichuan Li (2026). Measuring Whether LLM Tutors Teach or Solve: A Diagnostic for Educational Impact. arXiv:2606.16206. arXiv preprint.