Research Article
Comprehensive Review of Intelligent Tutoring Systems
Synthesis: Comprehensive Review of Intelligent Tutoring Systems — Journal of Computers in Education (2025). A systematic literature review covering 2010–2025 that analyzes the deployment and effectiveness of Intelligent Tutoring Systems (ITS) in real educational settings. The review examines the full landscape of ITS research — pedagogical strategies, natural language processing, adaptive learning mechanisms, student modeling approaches, and domain-specific applications — and arrives at a nuanced verdict: the evidence for ITS effectiveness is mixed, revealing a complex landscape of genuine advancements alongside persistent challenges in scientific rigor and real-world impact.
Key Findings
This review provides the most comprehensive mapping of the Intelligent Tutoring field since the emergence of Large Language Models (LLMs)-based tutoring approaches transformed the landscape. Spanning a 15-year window (2010–2025), it captures both the pre-LLM era of traditional ITS and the post-LLM era that has fundamentally reshaped what is technically possible.
Mixed effectiveness evidence. The review's central finding is that ITS effectiveness is neither uniformly positive nor categorically negative. ITS can improve student performance by roughly 20% on average, yet individual human tutoring still demonstrates up to 98% improvement — the cost/scalability gap ITS aim to fill. Some systems demonstrate substantial learning-gains, particularly in well-structured domains like mathematics and programming where Learner Modeling and Adaptive Instruction and Knowledge Tracing techniques are most mature. Other deployments show negligible or context-dependent effects. This mixed picture challenges both the optimistic narrative that AI tutoring is a proven solution and the pessimistic narrative that it is ineffective. Instead, it calls for more nuanced questions: which systems, for which learners, in which contexts, produce which outcomes? This aligns with the systematic-review literature's emphasis on contextual factors.
Pedagogical strategies. The review catalogs the range of pedagogical approaches embedded in ITS, from Socratic Method and Scaffolding to Adaptive Learning and Adaptive Learning pathways. A key finding is that many ITS implementations lack explicit pedagogical grounding — the tutoring behavior is often driven by technical capabilities (what the system can do) rather than pedagogical principles (what the system should do). This echoes concerns in the Training Pedagogical LLMs for Tutoring literature about the gap between technical sophistication and pedagogical intentionality.
NLP and adaptive mechanisms. The integration of Educational NLP techniques — including Automated Question Generation, short-answer assessment, and dialogue management — has advanced substantially over the review period. However, the review notes that many NLP components are evaluated in isolation rather than as integrated parts of tutoring systems that actually interact with learners. Similarly, Adaptive Learning show promise but often rely on narrow student models that fail to capture the full complexity of learner cognition and affect — a gap that the Affective Tutoring and Multimodal Dialogue in STEM Education communities are beginning to address.
Student modeling challenges. Learner Modeling and Adaptive Instruction remains both the foundation and the bottleneck for ITS. While Knowledge Tracing techniques (including Bayesian approaches like StanBKT: Rethinking Parameter Estimation in Bayesian Knowledge Tracing and deep learning variants) have improved, the review identifies persistent gaps in modeling higher-order cognitive processes, metacognition, and motivational states. This connects to the Engagement Intensity as a Learner-Modeling Signal for Adaptive AI Ethics Instruction and Metacognition literatures.
Scientific rigor deficit. One of the review's most important contributions is its methodological critique. Many ITS studies suffer from weak experimental designs — small sample sizes, absence of control groups, short intervention durations, and inadequate statistical analyses. The authors call for greater scientific rigor, including RCT where feasible, pre-registration of study designs, and transparent reporting aligned with educational research standards. This methodological critique connects to broader concerns in AI Ed Evaluation about the quality of evidence in AI education research.
Synthesis with Current Knowledge Base Evidence
| Claim in review | Supporting evidence in knowledge base | Contradictory evidence |
|---|---|---|
| ITS show mixed real-world effectiveness | The Evidence Base on AI in K-12: A 2026 Review (only 20/818 papers meet causal standards) | EduQwen (96.52% benchmark, but benchmark ≠ classroom) |
| Need for stronger experimental rigor | Hardy & Kim (benchmark≠teaching quality) | — |
| NLP advances for dialogue | Interpretable Knowledge Tracing (interpretable dialogue modeling) | SafeTutors (multi-turn degradation: 17.7% → 77.8%) |
| Affective computing as advancement | MathBuddy (+23 points win rate) | SafeTutors: Benchmarking Pedagogical Safety in AI Tutoring Systems (emotional risks, parasocial dependency) |
| Multi-agent architectures | Evolution of AI in Education: Agentic Workflows (four paradigms), Human-in-the-Loop (MAIC) | ProPACT (effective but requires eye-tracking hardware) |
Key Implications for the Field
- Benchmarks are not enough. High benchmark scores (CDPK, DeepTutor) must be complemented by classroom RCTs measuring actual learning gains.
- Teacher integration is a bottleneck. Technical sophistication matters less than curriculum fit and teacher control — see Human-in-the-Loop.
- Long-term studies are essential. Most ITS research measures immediate outcomes; SRL, metacognition, and transfer require longitudinal designs.
- Domain-specificity is real. A system effective in math may fail in writing; claiming "general tutoring" without domain evidence is overstated.
- Ethical and equity dimensions matter. Data privacy, algorithmic bias, and academic integrity are core determinants of whether ITS gains are sustainable and fair.
What this means for practice
- Researchers. State the tutoring principles a system is meant to enact before evaluating it, because the review finds many implementations driven by technical capability rather than pedagogical grounding, which leaves their behavior untraceable to any stated theory.
- Design RCT-quality studies that separate the effect of a specific ITS feature from confounds such as novelty, instructor quality, and student self-selection, and run them in real educational settings rather than laboratories.
- Pre-register the design and report a unified set of key performance indicators, since evaluation methods ranging from user surveys to pre/post testing currently block comparison and reproducibility across studies.
- Disaggregate outcomes by gender, socioeconomic background, and prior-knowledge level; the review finds demographic disaggregation consistently absent from reported results.
- Designers. Model metacognition, motivation, and affect alongside domain knowledge and embed Learning Analytics from the first deployment, then treat benchmark scores as a starting point rather than evidence of classroom effectiveness and plan explicitly for curriculum alignment, teacher training, and LMS integration.
Limitations
- Inclusion required a sample of at least 100 participants and a validation period of at least 6 weeks, so smaller but well-controlled studies are systematically excluded from the 127 articles reviewed.
- The review screened 37,617 records down to 127 articles plus 26 web reports, restricted to English-language, peer-reviewed work indexed in Web of Science, Scopus, IEEE Xplore, and Springer.
- Its effectiveness conclusions rest on studies the review itself judges to be mostly short-term and controlled, measured in part by self-reported engagement and satisfaction that may not correlate with measured learning.
- The headline figures — roughly 20% improvement for ITS and up to 98% for individual human tutoring — are drawn from heterogeneous evaluations with inconsistent outcome measures, so they are not a common effect size.
Citation
Zerkouk, M., Mihoubi, M., & Chikhaoui, B. (2025). Comprehensive Review of Intelligent Tutoring Systems. Journal of Computers in Education.