AI Ed Wiki logoAI Ed WikiUse with AI

Engagement, not model capability, is the binding constraint on AI tutoring. In a two-year cluster randomized trial across 18 Tennessee middle schools, randomly assigned students used Khan Academy with its AI tutor Khanmigo β€” configured to coach rather than give answers (a scaffolded guided-help design) β€” during existing daily remedial math sessions. Assignment raised math achievement by 1.3 national percentile ranks per term (~0.06–0.08 SD per school year; an implied ~0.14 SD for a full year of active participation), gains resembling those of Khan Academy practice without AI assistance. The reason: 96% of students tried Khanmigo at least once, but the median student messaged it on only about a third of practice days and in only 17% of the exercise sessions in which they made a mistake β€” and most messages were bare answers or clicks on suggested prompts, with just ~14.5% containing a genuine mathematical question or reasoning step. "Making AI effective for learning appears to be as much a behavioral challenge as a technological one": realized Help Seeking, not access, is what limits LLM tutoring in routine school settings.

Key Findings

  • Modest achievement gains at very low cost. Assignment raised math achievement ~0.06–0.08 SD per school year (1.3 national percentile ranks per term), delivered inside remedial blocks the district already staffed at a marginal license cost of roughly $15 per student per year β€” a cost structure that, unlike human tutoring, scales without adding skilled labor.
  • The effect is primarily the structured practice, not the tutor. Gains resembled those from Khan Academy practice without AI assistance; estimates are best interpreted as the effect of structured, individually targeted practice, with the tutor's contribution bounded by its low realized use. High-dosage human tutoring still delivers far more per student (0.2–0.4 SD) β€” software does not close that gap.
  • The AI tutor was barely engaged. 96% of students tried Khanmigo at least once, but the median student messaged it on only ~14% of exercise sessions and ~1/3 of practice days, and roughly one message in seven contained a genuine math question or reasoning step. The binding constraint was engagement, not what the tutor could do.
  • Struggling students are least likely to seek help unprompted. Consistent with the economics-of-education literature, interventions that rely on student initiative reach fewest of the students who would benefit most β€” especially when the action is incremental (considering one mistake on one problem). The limited dialogue reflects a lack of interest in the marginal act, not a defect of the AI tutor.
  • Design responds to the engagement margin. Khan Academy redesigned Khanmigo in 2026 to activate automatically during practice, citing lower student-initiated use than anticipated; evidence from other platforms indicates known human support raises engagement with AI tutoring where availability alone does not. Returns to AI-tutoring investment depend on the engagement margin as much as on capability.

Practical Implications

  • Projections should rest on realized use, not capability. The single most transferable lesson is that embedding a capable tutor in mandatory, teacher-supported session time produced substantial practice but thin tutoring dialogue β€” help-seeking remained a choice most students declined most of the time, consistent with behavioral barriers such as present bias, reliance on routine, and help avoidance.
  • Expect the tutor to sit idle unless it is made structurally salient. The 2026 Khanmigo redesign (auto-activating during practice) and cross-platform evidence that known human support raises engagement together point to the same design lever: reduce the marginal cost of initiating a tutoring exchange rather than only improving the tutor's responses.
  • The engagement margin is the priority for cost-effectiveness. At ~$15 per student per year the program already delivers a meaningful share of the individualized-feedback margin that historically required expensive human attention; raising the realized dosage of substantive dialogue is the highest-leverage improvement available.

Connected Concepts

Connected Articles

Citation

Oreopoulos, P., & Low, N. (2026). One click away: AI tutoring with Khanmigo in a two-year school experiment (NBER Working Paper No. 35620). National Bureau of Economic Research.