📄 Research Article
AI Tutor Effectiveness Review
Zerkouk, Mihoubi & Chikhaoui (2025) systematically analyzed qualified studies from 2010–2025 across:
A comprehensive systematic review of AI-based Intelligent Tutoring Systems (2010–2025) reveals a field with transformative potential but mixed real-world effectiveness, persistent implementation challenges, and a critical need for stronger experimental rigor.^Zerkouk Comprehensive Review ITS 2025
Review Scope
Zerkouk, Mihoubi & Chikhaoui (2025) systematically analyzed qualified studies from 2010–2025 across:
Key Findings
1. Mixed Effectiveness
Despite decades of progress and significant transformative potential, ITS have produced mixed results in real-world educational contexts. Effectiveness varies dramatically by:
2. Complex Advancement Landscape
The field shows both notable advancements and persistent challenges:
3. Scientific Rigor Gap
The review identifies a critical need for stronger experimental design and data analysis:
Synthesis with Current Wiki Evidence
| Claim in review | Supporting evidence in wiki | Contradictory evidence | |
|---|---|---|---|
| ITS show mixed real-world effectiveness | Stanford Evidence Base AI K12 2026 (only 20/818 papers meet causal standards) | [[pedagogical-llm-training | EduQwen]] (96.52% benchmark, but benchmark ≠ classroom) |
| Need for stronger experimental rigor | [[educational-llm-alignment | Hardy & Kim]] (benchmark≠teaching quality) | — |
| NLP advances for dialogue | Knowledge Tracing IRT (interpretable dialogue modeling) | [[ai-tutor-safety-harms | SafeTutors]] (multi-turn degradation: 17.7% → 77.8%) |
| Affective computing as advancement | [[affective-tutoring | MathBuddy]] (+23 points win rate) | AI Tutor Safety Harms (emotional risks, parasocial dependency) |
| Multi-agent architectures | Agentic Workflows Education (four paradigms), Human In The Loop AI (MAIC) | [[collaborative-ai-tutoring | ProPACT]] (effective but requires eye-tracking hardware) |
Implications for the Field
1. Benchmarks are not enough. High benchmark scores (CDPK, DeepTutor) must be complemented by classroom RCTs measuring actual learning gains.
2. Teacher integration is a bottleneck. Technical sophistication matters less than curriculum fit and teacher control — see Human In The Loop AI.
3. Long-term studies are essential. Most ITS research measures immediate outcomes; SRL, metacognition, and transfer require longitudinal designs.
4. Domain-specificity is real. A system effective in math may fail in writing; claiming "general tutoring" without domain evidence is overstated.
Connected Concepts
Connected Articles
Citation
Zerkouk, Mihoubi & Chikhaoui (2025). AI Tutor Effectiveness Review.