On this page

AI tutoring is not a monolith — a Stanford SCALE / National Student Support Accelerator (NSSA) policy brief (August 2026) mapping AI tutoring models along a "relational intensity" spectrum and weighing them against the evidence base for high-impact tutoring. The brief's core message: high-impact tutoring remains defined by live human-led instruction with strong student-tutor relationships; AI is best used to enhance tutor effectiveness and educator capacity, not to replace human-led high-impact tutoring. AI-led software offers potential for supplemental practice but does not yet meet the evidence base or definition of high-impact tutoring — and its effectiveness depends as much on integration (teachers in classrooms, parents at home) as on software quality.

Key Findings

  • AI tutoring is a spectrum, not a single thing. The brief maps models by relational intensity (the depth and consistency of human connection between student and tutor): human-led (in-person/remote tutoring, human tutoring with AI support) → AI-led with human support → AI-only tutoring. As direct human relationships decrease, the evidence base becomes thinner.
  • High-impact tutoring is well-evidenced; AI-led is not (yet). Students in high-impact tutoring show learning gains of three to more than 15 months across grade levels and content areas. The emergent AI-tutoring research base lags field implementation. A pair of RCTs across two school districts found human check-ins raised elementary students' engagement with an AI platform, but usage still fell short of the dosage associated with gains and reading achievement did not improve.
  • Dosage is the key driver — and AI struggles to meet it. AI-led tutoring inherits the dosage evidence base (e.g., ~90-minute weekly threshold) only when scheduled, supervised, and protected within the school day. In a study of 181,000 students on a supplemental math platform, only 5% reached the recommended 30 minutes/week and 41% never logged on; teacher/school/district factors explained 57% of usage variance. Even AI-only tutoring studies saw 40–47% of students never use the platform.
  • The student-tutor relationship is central and not replicable by AI. Relationship-building strategies — especially consistent tutoring with the same tutor — improve engagement, attendance, motivation, and learning outcomes. AI does not yet replicate human relationships; developers describe using AI strengths rather than replacing relationships.
  • Human oversight improves engagement and alignment but not dosage. Human supervision of AI-led sessions preserved structured learning goals and curriculum alignment, but did not by itself get students to the dosage that produces gains.
  • Safety and privacy are essential Guardrails. Safe direct-to-student AI requires strict student data privacy safeguards (enterprise-grade platforms, formal data-protection agreements, PII training) and attention to student-safety guardrails and the depth of unmonitored interaction. Open questions remain about AI companions' effects on developing minds and prosocial development.

Implications

  • Use AI to augment, not replace. Prioritize applications that expand human-led instruction and educator capacity — e.g., AI-generated curriculum-aligned practice materials, lesson preparation, real-time tutor support, scheduling, and data analysis.
  • Evaluate AI tutoring against the high-impact-tutoring elements. For any AI model, ask which conditions of effective tutoring it reproduces, changes, or drops (regularity, small-group/1:1 ratios, well-trained tutors, data-driven instruction, vetted materials, strong relationships) — more useful than asking whether "AI tutoring works."
  • Design for integration and dosage. AI tutoring's effectiveness depends on being embedded in school-day structures with enforced dosage and oversight; AI reintroduces an opt-in requirement that live school-day tutoring removes.
  • Treat evidence gaps seriously. AI-led and AI-only models sit in yellow/red zones with significant evidence gaps on effectiveness, safety, and developmental impact — warranting careful, staged adoption.

Connected Concepts

Connected Articles

Citation

Turano, C., Pihl, V., Agnew, C., Ziegler, L., & Loeb, S. (2026). AI Tutoring is Not a Monolith: What We Actually Know. Stanford SCALE / National Student Support Accelerator (NSSA), Stanford University. https://scale.stanford.edu