📄 Research Article
Methodologies for Improving the Quality of AI Tutoring in K-12 Education
Synthesis: Udeshi et al. (2026), the team behind Khanmigo (Khan Academy's K-12 AI tutor, launched 2023), describe the metrics they use to measure AI tutoring quality and student engagement, along with the live experiments that have moved those metrics. Given that LLMs are opaque black boxes, they argue robust evaluation and live experimentation are essential. The paper highlights changes across models, prompting, personalization, and agents that improved tutoring outcomes. Accepted at AIED 2026, it connects to AI Tutoring, Intelligent Tutoring, and Research Methods AIED literatures.
Evaluation as the Engine of Improvement
Many AI tutors leverage large language models today. Because LLMs are opaque black boxes, robust evaluation and live experimentation to measure the impact of every change are essential. Khan Academy pioneered AI-powered tutoring for K-12 with the launch of Khanmigo (2023).
What They Measure and Change
The paper describes the metrics used to measure AI tutoring quality and student engagement, and the various experiments run to improve them. It highlights the changes that moved their metrics, including changes to models, prompting, personalization, and agents.
Position
This practitioner account from a major edtech platform grounds the AI Tutoring and Intelligent Tutoring literature in real, large-scale K-12 deployment evidence, complementing controlled Research Methods AIED research.
Connected Concepts
Connected Articles
Citation
Udeshi, T., Khazenzon, A., Khan, K., Breen, N., Corwin, R. J., DiGiano, C., Weatherholtz, K., & Zaluski, M. (2026). Methodologies for improving the quality of AI tutoring in K-12 education. In Artificial Intelligence in Education (AIED 2026), LNCS vol. 16582. Springer. arXiv:2608.11259.