On this page

Meta-analysis and systematic review — the family of evidence-synthesis methods researchers use to aggregate and appraise a body of studies, rather than run a single new experiment. A systematic review applies a transparent, reproducible protocol to search, screen, appraise, and synthesize the literature on a focused question; a meta-analysis goes further by statistically pooling effect sizes across eligible studies to produce a weighted summary estimate and to test moderators. In AI in education, these methods are central to establishing the evidence base for whether AI tools work, under what conditions, and for whom — and to exposing gaps, bias, and the field's methodological quality.(A meta-analysis of the effect of generative AI on productivity and learning in programming)(Comprehensive Review of Intelligent Tutoring Systems)

Questions to Consider

  • Imagine you read ten studies on whether AI tutoring works — two show big gains, three show none, five show small positive effects. How would you decide what to conclude? That tension is exactly what systematic reviews and meta-analyses are built to resolve.
  • A systematic review and a meta-analysis are often treated as the same thing, but the page distinguishes them: a review synthesizes per a documented protocol, while a meta-analysis statistically pools effect sizes. When would pooling be inappropriate or impossible, even if a careful review exists?
  • Systematic reviews sit at the top of the evidence hierarchy partly because they compensate for small samples, heterogeneous designs, and conflicting results across studies. Where have you seen a single dramatic study shape opinion even though the pooled evidence was far more mixed?
  • Meta-analyses produce a weighted summary estimate — a single number like 'an average effect of 0.125 standard deviations.' What does a pooled average hide about the conditions, learners, or contexts where the effect differs — and why does that matter for whether you'd act on it?
  • Both methods commit to a transparent, reproducible protocol (often PRISMA) precisely because the choices of what to search and include can bias the result. How much would you trust a review that didn't disclose its search and screening decisions?

Introduction

Systematic reviews and meta-analyses sit at the top of the traditional evidence hierarchy precisely because they synthesize many individual studies, compensating for the small samples, heterogeneous designs, and conflicting results that characterize any fast-moving applied field. In AI in education, where new tools and studies appear constantly, reviews play the crucial role of taking stock: mapping what has been studied, aggregating what is known, and flagging where evidence is thin or methodologically weak. They differ from a narrative or integrative literature review, which provides qualitative synthesis, in their commitment to a documented protocol and (for meta-analysis) statistical pooling.(The AI Literacy Heptagon: A Structured Approach to AI Literacy in Higher Education)

Systematic review vs. meta-analysis

Systematic review Meta-analysis
Core activity Search, screen, appraise, synthesize studies per a documented protocol Statistically pool effect sizes across eligible studies
Output A narrative/thematic synthesis and evidence map, often with PRISMA flow A pooled effect estimate with confidence intervals, plus moderator analysis
Statistical pooling Optional (many reviews are qualitative) Required
When used Mapping a fragmented literature, answering "what has been studied and what does it show?" When multiple comparable quantitative studies exist, answering "how large is the effect overall?"
Strength Transparent, reproducible scope and appraisal Increased power and precision; detects moderators and heterogeneity

Both follow PRISMA (Preferred Reporting Items for Systematic Reviews and Meta-Analyses) as the reporting standard, which documents the search, screening, and inclusion process for transparency and reproducibility. An integrative review may follow PRISMA principles for transparency while stopping short of statistical pooling.(A systematic review of AI-powered collaborative learning in higher education: Trends and outcomes from the last decade)(The AI Literacy Heptagon: A Structured Approach to AI Literacy in Higher Education)

Evidence-synthesis in AI in education

What reviews accomplish

Systematic reviews and meta-analyses in AI in education serve several distinct purposes:

Screening reliability can be construct-dependent: in a PRISMA-ScR scoping review of 123 studies of GenAI, offloading and learner agency, title/abstract agreement (Fleiss' κ = 0.7103) fell to κ = 0.4884 at full text, which the authors attribute to the interpretive boundary between broad GenAI relevance and substantive relevance to the mechanism (Wang et al. (2026)).

Examples from the knowledge base

  • Meta-analysis of GenAI and programming — pools evidence on the productivity-learning trade-off, finding significant productivity gains but no significant learning gain (g ≈ 0), illustrating meta-analysis's ability to separate short-term efficiency from durable learning.(A meta-analysis of the effect of generative AI on productivity and learning in programming)

  • Meta-analysis of GenAI in student learning from a TPACK lens — pools 71 studies (74 effect sizes) to a medium-to-large effect (g = 0.752) that decomposes unevenly: strong cognitive (g = 0.831) and affective (g = 0.729) gains but a small, non-significant behavioral-engagement effect (SMD = 0.057, p = 0.828) (Liu & Zhong, 2025).

  • Systematic review of AI in VET — first systematic review of 26 studies, documenting the Constructivism-in-name, behaviorist-in-practice gap and the absence of workplace studies.(Artificial intelligence in vocational education and training: A systematic review of educational purposes, theoretical conceptualizations, and empirical effectiveness)

  • Systematic review of GenAI in higher education — maps opportunities, challenges, and pedagogical innovations across a five-year window.

  • Comprehensive ITS review — a systematic review of intelligent tutoring systems with a focus on methodological rigor.

  • Systematic review of ChatGPT and critical/creative thinking — synthesizes evidence on whether Large Language Models (LLMs) use supports or undermines higher-order thinking.

  • Evidence base for AI in K-12 — reviews the strength of evidence for AI tutoring in schools.

  • Meta-analysis of AI literacy interventions — three-level meta-analysis of 59 studies (172 effects, 7,211 participants) estimating a large overall effect (g = 0.837) while showing that effectiveness varies by region and learning-outcome focus (knowledge-focused interventions outperformed those targeting skills, attitudes, or Ethics).

  • AI Literacy Heptagon — an integrative literature review following PRISMA principles, illustrating qualitative synthesis that stops short of meta-analysis.(The AI Literacy Heptagon: A Structured Approach to AI Literacy in Higher Education)

  • Meta-analysis of digital and AI technologies for foreign-language skills — Chen and Wei (2026) pool 40 experimental and quasi-experimental studies (3,367 participants) to a large effect (g = 0.962, 95% CI 0.765 to 1.159) with high heterogeneity (I2 = 85.178) and publication-bias checks that support the estimate (Orwin's fail-safe N = 5853; Egger's test p = 0.84465). The instructive part is the design moderator: quasi-experiments reported g = 1.019 against g = 0.474 for true experiments, and the non-linear duration and tool-count patterns rest on subgroups of one or two studies.

  • Meta-analysis of GenAI and learning outcomes in higher education — Jing and colleagues (2026) pool 35 studies and 175 effect sizes to g = 0.53 (95% CI [0.48, 0.64]), largest for professional skills (g = 0.72), then academic performance (g = 0.46) and emotional attitude (g = 0.44). Duration moderates every dimension, and the decline is the point: academic-performance effects fall from g = 5.07 for interventions of 0 to 4 weeks to g = 0.64 at 4 to 12 weeks and g = 0.01 beyond 12 weeks (Q = 35.19, p < 0.001). Discipline type and learning method were not significant moderators.

  • Meta-analysis of K-12 AI education and AI literacy — Yu and colleagues (2026) pool 16 studies (57 effect sizes, 3,837 students, all effects positive) to g = 0.892 (95% CI [0.548, 1.236]), with trim-and-fill adjustment leaving it at g = 0.952. All three tested moderators (publication source, year, school level) came back null, and the authors attribute the 94.68% heterogeneity to measurement: the studies operationalized AI Literacy as everything from AI knowledge and ethics to attitudes, self-efficacy, and career interest.

  • Systematic review of GenAI in programming education — Yalcin and colleagues (2026) PRISMA-screen 151 records to 46 studies, of which 26 give no usable description of instructional approach and 38 examine a single tool (ChatGPT alone appears in 35 of 46). The review maps tools, languages, and venues reliably but can code pedagogy for fewer than half its corpus, a clean demonstration that thin primary-study reporting caps what a synthesis can conclude.

  • Scoping review and evidence map of NLP in student evaluation of teaching — Eicher and da Silva (2026) score 421 studies on a technical axis and four value dimensions, coding actionability on a six-level scale from asserted usefulness to measured impact on decisions, and reach a usable output in 258 of 421 studies (61.3%) against intended-user evaluation in only 49 (11.6%). Turning a gap into a coded dimension lets the review name the 49.7-point boundary between "the model works" and "the feedback changes teaching" as its central finding rather than a caveat.

  • PRISMA 2020 review of GenAI in healthcare scenario learning — Neto and colleagues (2026) systematically searched five databases (9 Nov 2025) for peer-reviewed GenAI studies across scenario-, case-, problem-, and Simulation-based healthcare education, screening 1,151 records down to 23 included studies appraised with the Mixed Methods Appraisal Tool (MMAT). Their thematic synthesis surfaced six cross-cutting themes anchored on prompt design as instructional specification, and documented gaps in validation standardization, longitudinal/comparative designs, and efficiency quantification — a template for rigorous domain-specific GenAI systematic review.

  • Systematic review of language educators and GenAI — a PRISMA-aligned review of 23 SSCI-indexed empirical studies (December 2022–September 2024) on how language educators perceive, adopt, and learn to integrate GenAI, synthesized through the Aristotelian knowledge typology of episteme (theoretical understanding), techne (practical skill), and phronesis (practical wisdom). It documents cautious, selective adoption weighted toward behind-the-scenes preparation, persistent competency gaps across all three knowledge types, and only three structured PD interventions — illustrating how a review can expose both the evidence base and the field's methodological gaps (here, the scarcity of structured professional-development studies).

  • Systematic review of AI and service delivery in global higher education — Nyamboga (2026) synthesizes 155 studies across five domains and grades the certainty of each with adapted CASP and JBI-informed criteria, rating AI integration and infrastructure readiness high, leadership moderate, and ethical governance and sustainability low. The review also turns a quality lens on its own corpus: 117 studies (75.5%) reported successful implementation, 124 (80.0%) covered only short-term impact, and just 23 (14.8%) reported ethics, equity or governance failures. A synthesis can report how certain the evidence is, domain by domain, rather than only the direction it points.

  • Systematic review of RL in education — Riedmann, Schaper & Lugrin (2025) applied a PRISMA-standard protocol to synthesize 89 Reinforcement Learning studies (2000–2024). The review is an instructive case study in synthesis methods and their limits: it maps a rapidly growing but methodologically uneven field — over half of studies (n = 54) reported no statistical testing — and runs effect-size analysis on only 15 papers suitable for pooling, reporting intermediate-to-large effects (Cohen's d) from live evaluations. It also performed publication-bias checks (funnel plot, Egger's test, PET-PEESE) that found no significant bias but had limited power (n = 6), and it stops at thematic-plus-limited-quantitative synthesis rather than a full meta-analysis precisely because heterogeneous evaluation protocols prevented wider pooling — a concrete illustration of the heterogeneity and garbage-in/garbage-out limitations described below.

  • Scoping review of critical thinking in HCI research on AI — Inie and colleagues (2026) map 80 empirical papers and deliberately perform no quality assessment, so "claimed" influence means the conclusion each paper states. Only 23 papers (29%) define critical thinking, 46 (57%) cannot be assigned a theoretical position, and 49 (61%) measure the construct through self-report, which is what allows 50 papers (62%) to report a positive influence. When a construct is this undefined and its measures incommensurable, mapping the field is more informative than pooling it, and withholding appraisal shifts the claim from an effect to a description of what studies asserted.

  • Mapping review of AI integration in higher education (FACETS + SAMR) — a PRISMA 2020 review screening 959 records down to 22 intervention studies, illustrating how a coding framework (FACETS: Form, AI use case, Context, Education focus, Technology, SAMR) plus an evaluative lens (SAMR) maps a fragmented literature and grades depth of transformation. Most included studies sat at SAMR Substitution/Augmentation, showing mapping reviews can reveal an integration field that is broad but shallow — an alternative to effect pooling when the aim is describing a landscape rather than estimating an effect.

  • Systematic review of teacher intervention in K-12 AI-based instruction — Lee (2026) screened 1,565 records down to 29 studies with two independent raters at every stage (κ = 0.655 screening, κ = 0.647 eligibility) and an MMAT quality appraisal that removed one study. It is an instructive example of synthesis that deliberately stops short of pooling: because the independent effect of teacher intervention could not be separated from AI system design, instructional structure and classroom context in most of the included studies, the review reports conditional outcomes and an explanatory framework of process, strategy and effect rather than an effect size — the same limitation that separates a review from a meta-analysis.

  • Systematic review of ethical values and norms in AIED — Agarwal and colleagues (2026) screened 736 records across Web of Science, ERIC, IEEE CSDL, and ACM DL (plus backward snowballing) down to 25 included articles, consolidating the fragmented AIED ethics literature into six main ethical values (non-discrimination, data stewardship, human oversight, goodwill, explicability, educational aptness) and mapping ethical norms onto a stakeholder-by-value matrix. The review illustrates how a systematic protocol can synthesize a conceptually fragmented, largely non-empirical literature (only three of 25 articles were methodology papers or original research) and turn it into an actionable framework — here, a foundation for AI Governance and policy.

  • Three-level meta-analysis of GenAI in PBL/PjBL — Chen and colleagues (2026) pool 22 controlled studies and 66 effect sizes to a large effect on learner-internal outcomes (g = 0.819, 95% CI [0.655, 0.983]), then report that a PET-PEESE small-study correction lowers the estimate to g = 0.378 while the product-performance model (g = 1.958, six studies) carries a prediction interval that crosses zero. Two further studies were dropped for failing a What Works Clearinghouse baseline-equivalence threshold of 0.25 SD. The reported effect is contingent on the bias corrections and design screens a synthesis applies, and both belong in the interpretation rather than a footnote.

Interdisciplinary review of AI for dyslexia

Dabaghi, D'Urso & Sciarrone (2026) present a PRISMA-guided, interdisciplinary systematic review (2018–2024, n=72) of AI and generative AI to support students with dyslexia in education. The review maps AI across detection, assistive support, and personalized learning, finding these strands evolve in parallel rather than in integration, driven more by technological opportunity than by consolidated educational theory. It documents that generative AI is under-utilized in this domain (GAI research, all from 2024, clusters into chatbots, teacher-training support, and exploratory studies) and that ML-based help-education tools fall into five areas (specific applications, engagement, personalization, recommendation, generic support) while emphasizing technical performance over ecological validity. Open challenges include limited experimental validation, scalability and Accessibility of diagnostic tools, ethics/privacy concerns with sensitive student data, limited teacher support, and language/cultural barriers. The review's own methodological limitations — interpretative classification bias, exclusion of non-English studies, heterogeneous evaluation protocols that prevent quantitative synthesis, and a rapidly evolving GAI evidence base — illustrate the systematic-review family's core tension: a transparent protocol can map a fragmented field, but heterogeneous evaluation prevents statistical pooling, so the review stops at thematic synthesis rather than meta-analysis.

AI-era synthesis challenge: productivity vs. learning

Reviews of generative-AI interventions face a distinctive challenge that the knowledge base's synthesis research highlights: separating productivity gains from durable learning gains. Because generative AI can inflate immediate task performance (homework, assisted practice) without producing learning, meta-analyses must be careful about which outcome they pool. The GenAI-and-programming meta-analysis found large productivity gains but no significant learning gain (g ≈ 0) — a clean illustration. Large-scale field studies and unassisted-measure research show that the measured effect depends on whether outcomes are AI-assisted or proctored/unassisted. Reviews should therefore report assisted and unassisted outcomes separately, distinguish performance from learning, and flag studies that measure only immediate AI-supported performance. A complementary caution emerges from the AI-literacy meta-analysis: which outcome is pooled also shapes the answer — knowledge-focused AI Literacy interventions showed larger effects than those targeting skills, attitudes, or ethics, so a review that pools only knowledge outcomes can overstate what AI-literacy instruction achieves overall. This connects to AI Ed Evaluation and Summative Assessment.

The same artifact shows up in the time dimension. Jing and colleagues' 2026 meta-analysis of GenAI in higher education (35 studies, 175 effect sizes, pooled g = 0.53) reports academic-performance effects that fall from g = 5.07 for interventions of 0 to 4 weeks to g = 0.64 at 4 to 12 weeks and g = 0.01 beyond 12 weeks (Q = 35.19, p < 0.001), with emotional-attitude effects dropping from a peak of g = 2.05 to g = 0.18 over the same span. Effects that large, that early, are mostly registering task-level assistance rather than durable learning, so a pooled estimate dominated by short studies should be read as an upper bound on what a real implementation will sustain. Reviews that pool outcomes should therefore record intervention duration alongside the effect size, because duration is one of the few moderators that has held up across these syntheses.

Strengths and limitations

Strengths:

  • Efficient synthesis of a large, fragmented literature
  • Meta-analysis yields pooled effect estimates, increases statistical power, and detects moderators and heterogeneity
  • Systematic protocols improve transparency and reproducibility over narrative reviews
  • Essential for evidence-based practice and for identifying research gaps

Limitations:

Meta-research warns the AIED synthesis base is currently weak. A growing set of critiques documents that the field's headline AI-effect sizes — especially from early meta-analyses — are inflated by publication bias, construct incoherence, and methodological shortcuts. Bartoš et al. (2026), meta-analyzing 1,840 effect sizes from 67 meta-analyses, estimate the publication-bias-adjusted AI effect at roughly one-third the reported magnitude (SMD ≈ 0.196), with extreme heterogeneity. O'Neill (2026) audits 14 high-impact AIED meta-analyses and finds none had a coherent construct, valid publication-bias assessment, or resolved heterogeneity; twelve treated dependent effect sizes as independent. Weidlich et al. (2025) show most primary comparisons lack a well-defined treatment, control, and learning measure. This means readers should treat pooled AIED effect sizes as upper bounds until synthesis quality improves — see Limitations in AIEd Research for the full analysis.

  • Reporting standards have not kept pace with automation. PRISMA-LLM analyses SciLitBench, a corpus of 888 review-automation papers with 14,726 annotations, and documents a growth rate near 4.7% per month alongside an accountability gap: since 2023, 38.0% of software and product papers reported no evaluation at all, against 9.3% of LLM papers, and 52% of positive-only LLM evaluations reported an unmet high-bar concern. Because LLM and software pipelines now participate in stages that can alter the evidence base, the framework requires disclosure of where in the review workflow automation operated, what was evaluated, and which limitations were checked — a direct extension of the transparency problem that AIED review critiques have documented for meta-analyses of learning effects. (PRISMA-LLM: An Empirical Reporting Framework for AI-Assisted Systematic Reviews)

Relationship to other methods

Within the knowledge base's methodological landscape, meta-analysis and systematic review are the synthesis family, complementing primary designs:

  • Primary studies (experiments, surveys, qualitative work, design-based research) generate individual findings; reviews aggregate them. See Research Methods in AIED.
  • Effect-size reporting in primary studies (e.g., RCTs) is what makes later meta-analysis possible — reviews depend on studies reporting comparable, extractable effect sizes.
  • Evaluation (AI Ed Evaluation, Benchmark) assesses individual systems; reviews assess the literature on systems and interventions.
  • Educational measurement (Educational Measurement, Assessment Validity) concerns the quality of the outcome measures that reviews pool.

Implications for researchers

  1. Report extractable effect sizes. For a literature to be meta-analyzable, primary studies must report comparable effect sizes and adequate methods detail — a responsibility of every AIED study.(Research Methods in AIED) Adequate detail includes the instructional approach: a systematic review of GenAI in programming education could code pedagogy for fewer than half its 46 studies because 26 never described what they did, and the detail primary authors omit is the detail later syntheses cannot recover.
  2. Follow a transparent protocol. PRISMA-guided search, screening, and appraisal make reviews reproducible and defensible.
  3. Interpret pooled effects cautiously. Attend to heterogeneity, publication bias, and the quality of included studies before drawing strong conclusions.
  4. Use reviews to set the agenda. Reviews' documented gaps (failure cases, workplace settings, non-English and non-indexed work, long-term outcomes) should guide where new primary research is needed.(Artificial intelligence in vocational education and training: A systematic review of educational purposes, theoretical conceptualizations, and empirical effectiveness)
  5. Treat null moderators as findings. When moderators come back null and heterogeneity stays high, that is evidence about the state of the field (construct incoherence, thin design reporting) as much as about the intervention, and it belongs in the synthesis narrative rather than being dropped.(The Effects of K-12 Artificial Intelligence Education in Enhancing AI Literacy: A Meta-Analysis)(The Impact of Digital and Artificial Intelligence Technologies on the Improvement of Foreign Language Listening, Speaking, Reading and Writing Skills: A Meta-Analysis)

GenAI in Healthcare Scenario Learning

  • PRISMA 2020 review of GenAI in healthcare scenario learning. Neto and colleagues (2026) systematically searched five databases (9 Nov 2025) for peer-reviewed GenAI studies across scenario-, case-, problem-, and Simulation-based healthcare education, screening 1,151 records down to 23 included studies appraised with the Mixed Methods Appraisal Tool (MMAT). Their thematic synthesis surfaced six cross-cutting themes anchored on prompt design as instructional specification, and documented gaps in validation standardization, longitudinal/comparative designs, and efficiency quantification — a template for rigorous domain-specific GenAI systematic review.

Connected Concepts

Connected Articles

Connected FAQs

Embed this page

Copy the code below to embed a chromeless version of this page in a learning management system or other website. The embedded view hides the site header, navigation, and footer.