On this page

Synthesis: Generative AI technologies and educational outcomes: a comprehensive meta-analysis comparing traditional and AI-driven approaches — Dong (2026), Humanities and Social Sciences Communications 13, 559. A PRISMA-based meta-analysis of 53 studies that pools the effect of generative AI (GenAI) on educational outcomes. The central conclusion: GenAI generally outperforms traditional and non-GenAI approaches on academic achievement, higher-order thinking, and writing skills, and GenAI feedback is particularly effective — though game-assisted GenAI shows no significant added benefit, and gains vary by country and are consistent across university and secondary levels.

Key Findings

GenAI broadly improves learning outcomes. Pooling 53 studies, the meta-analysis finds GenAI-assisted education significantly outperforms non-GenAI approaches across academic achievement (g = 0.40), higher-order thinking (g = 0.72), learning motivation (g = 0.81), and writing skills (g = 0.76). This positions GenAI as a genuine Learning Gains lever rather than a novelty. The authors frame GenAI as a "cognitive aid" that strengthens thinking when used to support — not replace — the learner, echoing the Constructivism emphasis on active knowledge construction over passive reception. The pattern is consistent with broader AI Ed Evaluation syntheses that find positive but heterogeneous effects of GenAI in education.

GenAI feedback is the standout mechanism. GenAI feedback produced the largest effect in the study (g = 1.27), outperforming non-GenAI feedback. Its advantages are comprehension, timeliness, and objectivity in a context where human teachers lack time for detailed personalized feedback. This connects directly to the AI Feedback Quality literature. Caveats remain: students can distrust GenAI feedback, and its lack of emotional response may raise cognitive load. The authors ground the mechanism in Metacognition theory — feedback supports students' monitoring and self-regulation of learning.

Game-assisted GenAI shows no added value. Game-assisted GenAI did not significantly outperform non-game-assisted approaches (g = 0.24), because games can distract learners and reduce retention. The authors argue game design quality — not Gamification per se — determines effectiveness, in line with Self-Determination Theory. For younger learners with weaker self-AI Regulation in Education, non-gamed designs may be preferable.

Country moderators are nuanced. GenAI significantly outperformed non-GenAI in China (g = 0.71) and Pakistan (g = 0.75), where GenAI helps bridge educational-resource inequity. But in Korea and Turkey (g = 1.68, g = 0.01, both non-significant), teacher-centered pedagogical cultures and curricular misalignment slowed adoption. The authors caution that excluding Korean/Turkish-language studies may have biased these findings.

Consistent gains across educational levels. GenAI improved outcomes at both university (g = 0.70) and secondary (g = 0.80) levels, supporting its use across Higher Education and K-12 contexts. Strategies differ: university settings emphasize research skills and scholarly work; secondary settings emphasize knowledge delivery and interest cultivation.

Robustness. Publication bias was detected (Egger's and Begg's tests) but trim-and-fill analysis confirmed the pooled result remains significant. Leave-one-out sensitivity analysis showed all effect sizes within pooled confidence intervals, indicating stable, reliable findings.

Connection to the broader knowledge base

For AI Ed Evaluation practice, this study is a high-confidence, quantitative anchor showing that GenAI's benefits are real but context-dependent. The strong feedback finding reinforces the growing AI Feedback Quality evidence base and suggests institutions should prioritize feedback-intensive GenAI uses. The null game-assisted result is a useful counterweight to Game-Based Learning hype, and the country-level divergence warns against universalist claims: efficacy hinges on pedagogical culture and resource context. The authors recommend future research on students with disabilities and on long-term effects on self-directed learning and Critical Thinking — concerns that connect to the Over-Reliance literature on GenAI-induced cognitive erosion. Because the paper is a Meta-Analysis and Systematic Review, its pooled estimates help ground Generative AI adoption decisions in aggregated evidence rather than single studies.

What this means for practice

  • Instructors. Put generative AI where feedback is otherwise unaffordable: the pooled effect for GenAI feedback (g = 1.27) is the largest in the study and far above the achievement gain (g = 0.40), with the authors crediting its comprehension, timeliness and objectivity. Carry their two caveats into the classroom — students can distrust AI feedback, and its lack of emotional response may raise cognitive load.
  • Instructors. Expect the largest movement on higher-order thinking (g = 0.72), motivation (g = 0.81) and writing (g = 0.76), and use the tool to support the learner's own thinking rather than replace it, which is the "cognitive aid" framing the authors adopt and a direct echo of the over-reliance caution.
  • Designers. Do not assume engagement features earn their place: game-assisted generative AI showed no significant added benefit in this pooling, which is a useful counterweight to the assumption that gamified delivery improves outcomes by itself.
  • Administrators. Treat g = 0.40 as an average with real spread rather than a promise: the subgroup analyses diverge by country, so local evidence still decides, and the pooled estimate is best used as a prior to test against your own students rather than a guarantee of gain.

Limitations

  • The review may not include every relevant study: the authors name limits on their library access as a possible source of bias in the conclusions, so the pooled estimates rest on the literature they could reach.
  • The included studies were heterogeneous, which the authors say may have affected the meta-analytic results; the effect sizes are averages across unlike designs, dosages and populations rather than estimates for any one of them.
  • The technology moved during the study window: the authors note that generative AI capabilities were developing rapidly throughout, so their pooled effects describe the tools studied rather than whatever version a reader has in front of them now.
  • Pooled effects sit consistently across university and secondary levels, but the country-level subgroups diverge, so a single national or institutional setting is not what these estimates describe.

Citation

Dong, Y. (2026). Generative AI technologies and educational outcomes: a comprehensive meta-analysis comparing traditional and AI-driven approaches. Humanities and Social Sciences Communications 13, 559. https://doi.org/10.1057/s41599-026-06903-y

Embed this page

Copy the code below to embed a chromeless version of this page in a learning management system or other website. The embedded view hides the site header, navigation, and footer.