Research Article
The Impact of Artificial Intelligence-Supported Instruction on Student Learning in STEM: A Systematic Review and Meta-Analysis
Synthesis: This systematic review and meta-analysis pools 35 experimental and quasi-experimental studies published between 2005 and 2025 to estimate how AI-supported instruction affects student learning in STEM education. Using Hedges' g under a random effects model, the authors report a moderate to strong positive overall effect (g = 0.670, 95% CI [0.491, 0.848], p < 0.001), largest at the high school level and for interventions lasting more than one month and up to two months. Differences between science, mathematics and technology/engineering were not statistically significant, and the publication bias diagnostics suggested the overall result was robust, although heterogeneity across studies was substantial and the authors applied no formal quality appraisal tool.
Key Findings
- Across the 35 included studies, the random effects estimate of AI-supported instruction on STEM achievement was g = 0.670 (95% CI [0.491, 0.848], Z = 7.35, p < 0.001), described by the authors as a moderate to strong positive effect relative to traditional teaching.
- The fixed-effect counterpart for the same 35 studies was g = 0.593 (95% CI [0.522, 0.664], Z = 16.36, p < 0.001). The authors prefer the random effects model because it incorporates between-study variance into the estimate and is therefore more conservative and generalizable (effect-size synthesis).
- Heterogeneity must be read model by model, as the paper reports it separately: the fixed-effect analysis shows a high level of inconsistency (Q = 199.76, I2 = 82.98%), while the random effects analysis reports I2 = 15.75% once between-study variance (τ2) is modeled. The high fixed-effect heterogeneity is the authors' main reason for not treating AI as uniformly effective.
- Educational level was the strongest and only clearly significant moderator (between-group Q = 30.13, df = 3, p < 0.001): high school showed the largest effect (g = 1.099, 95% CI [0.89, 1.30], Z = 10.53, p < 0.001), followed by university (g = 0.578, 95% CI [0.48, 0.68], Z = 11.67, p < 0.001), elementary or primary (g = 0.465) and middle school (g = 0.392).
- Intervention duration also moderated effects (between-group Q = 16.95, df = 5, p = 0.004). The strongest band was more than one month and up to two months (g = 0.833, 95% CI [0.67, 0.99], Z = 10.30, p < 0.001); short interventions of up to five hours were also significant (g = 0.621, p = 0.003); the weakest band, more than one and up to seven days, was not statistically significant (g = 0.256, p = 0.070).
- There was no consistent pattern of increasing effectiveness with longer duration. Duration bands of more than two and up to three months (g = 0.618) and more than three months (g = 0.576) sat below the one-to-two-month band, which the authors read as evidence that instructional quality matters more than mere exposure time.
- Subject area differences were small and not statistically significant (between-group Q = 4.85, df = 2, p = 0.088): science g = 0.676, mathematics g = 0.650, technology and engineering g = 0.501. The authors explicitly caution that cross-disciplinary differences should be interpreted with caution and instead read the result as cross-disciplinary potential for adapted tools.
- Class size produced a marginally significant between-group difference (Q = 4.61, df = 2, p = 0.099), with the largest effect in medium-sized classes (g = 0.687), then large classes (g = 0.545) and small classes (g = 0.448).
- Publication bias diagnostics pointed to minimal distortion. The funnel plot was generally symmetrical around the mean effect with a few high-effect outliers on the right; Egger's regression intercept test was not significant (p = 0.095), while Begg's test was borderline (tau = −0.2303, p = 0.0517), which the authors ask readers to weigh with caution.
- Trim-and-fill imputed only one missing study and adjusted the mean effect upward from g_mean = 0.6799 to 0.7269 (95% CI [0.6696, 0.7841]). The increase rather than decrease is attributed to sampling variability, and the authors conclude publication bias did not substantially change the conclusions.
- Rosenthal's fail-safe number was 2404, meaning that many null-effect studies would be needed to make the pooled result non-significant; by Rosenthal's (1979) criterion the authors treat the meta-analytic findings as robust.
- Theoretical interpretation draws on cognitive load theory, Constructivism and sociocultural accounts of AI as a mediational artifact within the zone of proximal development, and Technological Pedagogical Content Knowledge (TPACK) and SAMR; the authors conclude AI-supported instruction mostly augments or modifies existing practice rather than transforming it.
- The overall picture is conditional rather than universally positive: AI is not inherently effective but becomes effective under specific pedagogical and contextual conditions, with learner readiness, design quality and implementation driving the variation.
Study Design and Method
The review followed PRISMA 2020 guidance and the protocol was registered on OSF. Searches ran in September 2025 across Web of Science, Scopus, ERIC, ScienceDirect, Google Scholar and the Council of Higher Education National Thesis Center, in Turkish and English, using Boolean strings combining AI terms (chatbot, intelligent tutoring system, adaptive learning, machine learning, deep learning) with STEM and achievement terms. The window opened in January 2005 and closed on 1 September 2025.
The PRISMA flow records 450 records identified, 86 duplicates removed before screening, 364 titles and abstracts screened with 319 excluded, 45 reports sought for retrieval with 2 not retrievable, and 43 full-text reports assessed. Of these, 8 were excluded for insufficient quantitative data (n = 5) or poor methodological quality (n = 3), leaving 35 studies. Most included studies (n = 29) were published after 2020.
Inclusion required an AI-supported intervention in a STEM field, an experimental or quasi-experimental design with comparison groups, measurable achievement outcomes and sufficient statistics (N, M, SD) for an effect size; book chapters and letters to the editor were excluded, while preprints, theses and conference papers were considered. Three researchers coded the studies: they first coded pilot studies jointly to standardize the form, then randomly assigned studies to two independent coders, reaching Cohen's Kappa κ = 0.93. Only one effect size was extracted per study to preserve statistical independence, and Hedges' g was chosen over Cohen's d because the correction factor reduces small-sample bias. Analyses used Comprehensive Meta-Analysis (CMA) 3.0.
Implications
The authors argue AI should be treated as a pedagogical partner rather than an autonomous instructional agent. Because effectiveness depends on design quality, they call for AI tools aligned with learning objectives and structured to manage cognitive load, and warn that poorly designed implementations risk superficial engagement, over-reliance on automated feedback, or overload. They stress teacher competence to evaluate AI outputs critically, and flag equity risk: differences in infrastructure, learner readiness and resources could let AI reinforce existing inequalities, so they recommend sustained, system-level integration instead of isolated short-term pilots.
Limitations
The authors state several limitations directly. Heterogeneity remained high, leaving a substantial share of variance unexplained and pointing to unmeasured moderators such as implementation fidelity, teacher involvement and learner characteristics. The predominance of short- to medium-term interventions limits what can be said about long-term Sustainability and transfer, so longitudinal work is needed. The evidence base is mostly cognitive and achievement-focused, so affective, motivational and metacognitive outcomes are underrepresented, and the role of generative AI in relation to accuracy, trust and critical thinking needs more study. Cross-cultural and contextual comparisons remain underexplored. Critically, the authors did not apply a formal quality appraisal tool: methodological quality was operationalized through the predefined inclusion and exclusion criteria, which they describe as ensuring a minimum threshold of rigor, and the three studies excluded for poor quality were judged on that basis rather than a validated appraisal instrument. The borderline Begg's test (p = 0.0517) is also flagged as a reason for interpretive caution.
Connected Concepts
- Meta-Analysis and Systematic Review — the study's method: pooled effect sizes from experimental STEM studies
- STEM Education — the setting and population for all included interventions
- Learning Gains — the achievement outcomes synthesized (tests, exams, standardized assessments)
- Intelligent Tutoring — one of the main AI intervention types in the pooled studies
- Adaptive Learning — adaptive platforms and sequenced feedback as intervention designs
- Personalized Learning — the personalization mechanism the authors credit for the effect
- Science Education — highest-effect subject area subgroup
- Math Education — closely matched subject area subgroup
- Engineering Education — lower-effect technology and engineering subgroup
- Higher Education — the largest subgroup by number of studies (university level)
- K-12 — primary, middle and high school subgroups, where effects diverged sharply
- AI in Education — the broader field the synthesis contributes to
Connected Articles
- Effect of Artificial Intelligence on Learning: A Meta-Meta-Analysis — meta-meta-analysis of AI on learning outcomes
- Generative AI technologies and educational outcomes: a comprehensive meta-analysis comparing traditional and AI-driven approaches — meta-analysis of generative AI effects on educational outcomes
- Why does AI unlock new possibilities in STEM education? A Bibliometric Analysis of Trends and Future Agenda — bibliometric mapping of AI research in STEM
- Mapping the Scaffolding of Metacognition and Learning by AI Tools in STEM Classrooms: A Bibliometric-Systematic Review — review of AI and metacognition in STEM settings
- Artificial Intelligence in Science and Chemistry Education: A Systematic Review — systematic review of AI in science and chemistry education
- A meta-analysis of the effect of generative AI on productivity and learning in programming — meta-analysis of generative AI in programming learning
- Developing Deep Learning in Science Through an Adaptive AI-Based STEM Instructional Program: Evidence From Sixth-Grade Classrooms — adaptive AI in STEM with deep learning methods
- Methodologies for Improving the Quality of AI Tutoring in K-12 Education — methodological critique of K-12 AI tutoring evidence
Citation
Doğan, Y., Kılıç, Z., Kalınkara, Y., & Talan, T. (2026). The Impact of Artificial Intelligence-Supported Instruction on Student Learning in STEM: A Systematic Review and Meta-Analysis. Journal of Intelligence, 14(6), 109.