AI Tutor Effectiveness Review

Created: 2026-05-07 | Tags: intelligent-tutoringbenchmarkefficacy-studyhigher-edk-12
πŸ“„ Full text: arXiv:2507.18882 Β· local
A comprehensive systematic review of AI-based Intelligent Tutoring Systems (2010–2025) reveals a field with transformative potential but mixed real-world effectiveness, persistent implementation challenges, and a critical need for stronger experimental rigor.^zerkouk-comprehensive-review-its-2025

Review Scope

Zerkouk, Mihoubi & Chikhaoui (2025) systematically analyzed qualified studies from 2010–2025 across:

Key Findings

1. Mixed Effectiveness

Despite decades of progress and significant transformative potential, ITS have produced mixed results in real-world educational contexts. Effectiveness varies dramatically by:

2. Complex Advancement Landscape

The field shows both notable advancements and persistent challenges:

3. Scientific Rigor Gap

The review identifies a critical need for stronger experimental design and data analysis:

Synthesis with Current Wiki Evidence

Claim in review Supporting evidence in wiki Contradictory evidence
ITS show mixed real-world effectiveness ai-k12-evidence-base (only 20/818 papers meet causal standards) EduQwen (96.52% benchmark, but benchmark β‰  classroom)
Need for stronger experimental rigor Hardy & Kim (benchmark≠teaching quality) —
NLP advances for dialogue knowledge-tracing-irt (interpretable dialogue modeling) SafeTutors (multi-turn degradation: 17.7% β†’ 77.8%)
Affective computing as advancement MathBuddy (+23 points win rate) ai-tutor-safety-harms (emotional risks, parasocial dependency)
Multi-agent architectures agentic-workflows-education (four paradigms), human-in-the-loop-ai (MAIC) ProPACT (effective but requires eye-tracking hardware)

Implications for the Field

1. Benchmarks are not enough. High benchmark scores (CDPK, DeepTutor) must be complemented by classroom RCTs measuring actual learning gains. 2. Teacher integration is a bottleneck. Technical sophistication matters less than curriculum fit and teacher control β€” see human-in-the-loop-ai. 3. Long-term studies are essential. Most ITS research measures immediate outcomes; SRL, metacognition, and transfer require longitudinal designs. 4. Domain-specificity is real. A system effective in math may fail in writing; claiming "general tutoring" without domain evidence is overstated.

Related Pages