π Research Article
One Click Away: AI Tutoring with Khanmigo in a Two-Year School Experiment
Engagement, not model capability, is the binding constraint on AI tutoring. In a two-year cluster randomized trial across 18 Tennessee middle schools, randomly assigned students used Khan Academy with its AI tutor Khanmigo β configured to coach rather than give answers (a scaffolded guided-help design) β during existing daily remedial math sessions. Assignment raised math achievement by 1.3 national percentile ranks per term (~0.06β0.08 SD per school year; an implied ~0.14 SD for a full year of active participation), gains resembling those of Khan Academy practice without AI assistance. The reason: 96% of students tried Khanmigo at least once, but the median student messaged it on only about a third of practice days and in only 17% of the exercise sessions in which they made a mistake β and most messages were bare answers or clicks on suggested prompts, with just ~14.5% containing a genuine mathematical question or reasoning step. "Making AI effective for learning appears to be as much a behavioral challenge as a technological one": realized Help Seeking, not access, is what limits LLM tutoring in routine school settings.
Key Findings
- Modest achievement gains at very low cost. Assignment raised math achievement ~0.06β0.08 SD per school year (1.3 national percentile ranks per term), delivered inside remedial blocks the district already staffed at a marginal license cost of roughly $15 per student per year β a cost structure that, unlike human tutoring, scales without adding skilled labor.
- The effect is primarily the structured practice, not the tutor. Gains resembled those from Khan Academy practice without AI assistance; estimates are best interpreted as the effect of structured, individually targeted practice, with the tutor's contribution bounded by its low realized use. High-dosage human tutoring still delivers far more per student (0.2β0.4 SD) β software does not close that gap.
- The AI tutor was barely engaged. 96% of students tried Khanmigo at least once, but the median student messaged it on only ~14% of exercise sessions and ~1/3 of practice days, and roughly one message in seven contained a genuine math question or reasoning step. The binding constraint was engagement, not what the tutor could do.
- Struggling students are least likely to seek help unprompted. Consistent with the economics-of-education literature, interventions that rely on student initiative reach fewest of the students who would benefit most β especially when the action is incremental (considering one mistake on one problem). The limited dialogue reflects a lack of interest in the marginal act, not a defect of the AI tutor.
- Design responds to the engagement margin. Khan Academy redesigned Khanmigo in 2026 to activate automatically during practice, citing lower student-initiated use than anticipated; evidence from other platforms indicates known human support raises engagement with AI tutoring where availability alone does not. Returns to AI-tutoring investment depend on the engagement margin as much as on capability.
Practical Implications
- Projections should rest on realized use, not capability. The single most transferable lesson is that embedding a capable tutor in mandatory, teacher-supported session time produced substantial practice but thin tutoring dialogue β help-seeking remained a choice most students declined most of the time, consistent with behavioral barriers such as present bias, reliance on routine, and help avoidance.
- Expect the tutor to sit idle unless it is made structurally salient. The 2026 Khanmigo redesign (auto-activating during practice) and cross-platform evidence that known human support raises engagement together point to the same design lever: reduce the marginal cost of initiating a tutoring exchange rather than only improving the tutor's responses.
- The engagement margin is the priority for cost-effectiveness. At ~$15 per student per year the program already delivers a meaningful share of the individualized-feedback margin that historically required expensive human attention; raising the realized dosage of substantive dialogue is the highest-leverage improvement available.
Connected Concepts
- Intelligent Tutoring
- LLM
- Generative AI
- Help Seeking
- Student Engagement
- Edtech Platform
- Math Education
- K 12
- Adaptive Learning
- Learning Gains
- Equity In AI Education
Connected Articles
- Virtual Tutoring Computer Assisted Learning Takeup 2026 β Virtual tutoring with CAL: an experiment in take-up and learning
- Making AI Tutoring Productive Mastery Math 2026 β Making AI tutoring productive: mastery-based math practice
- Elevate GenAI Virtual Tutors β GenAI virtual tutors
- Access Not Enough AI Tutoring 2026 β Access is not enough for AI tutoring
- AI Tutoring Quality K12 Methodologies 2026 β Improving AI tutoring quality in K-12
Citation
Oreopoulos, P., & Low, N. (2026). One click away: AI tutoring with Khanmigo in a two-year school experiment (NBER Working Paper No. 35620). National Bureau of Economic Research.