AI Ed Wiki logoAI Ed WikiUse with AI

Meta-analysis and systematic review β€” the family of evidence-synthesis methods researchers use to aggregate and appraise a body of studies, rather than run a single new experiment. A systematic review applies a transparent, reproducible protocol to search, screen, appraise, and synthesize the literature on a focused question; a meta-analysis goes further by statistically pooling effect sizes across eligible studies to produce a weighted summary estimate and to test moderators. In AI in education, these methods are central to establishing the evidence base for whether AI tools work, under what conditions, and for whom β€” and to exposing gaps, bias, and the field's methodological quality.^GenAI Meta Analysis Programming Learning^Zerkouk Comprehensive Review ITS 2025

Systematic reviews and meta-analyses sit at the top of the traditional evidence hierarchy precisely because they synthesize many individual studies, compensating for the small samples, heterogeneous designs, and conflicting results that characterize any fast-moving applied field. In AI in education, where new tools and studies appear constantly, reviews play the crucial role of taking stock: mapping what has been studied, aggregating what is known, and flagging where evidence is thin or methodologically weak. They differ from a narrative or integrative literature review, which provides qualitative synthesis, in their commitment to a documented protocol and (for meta-analysis) statistical pooling.^AI Literacy Heptagon 2026

Systematic review vs. meta-analysis

Systematic reviewMeta-analysis
Core activitySearch, screen, appraise, synthesize studies per a documented protocolStatistically pool effect sizes across eligible studies
OutputA narrative/thematic synthesis and evidence map, often with PRISMA flowA pooled effect estimate with confidence intervals, plus moderator analysis
Statistical poolingOptional (many reviews are qualitative)Required
When usedMapping a fragmented literature, answering "what has been studied and what does it show?"When multiple comparable quantitative studies exist, answering "how large is the effect overall?"
StrengthTransparent, reproducible scope and appraisalIncreased power and precision; detects moderators and heterogeneity

Both follow PRISMA (Preferred Reporting Items for Systematic Reviews and Meta-Analyses) as the reporting standard, which documents the search, screening, and inclusion process for transparency and reproducibility. An integrative review may follow PRISMA principles for transparency while stopping short of statistical pooling.^AI Collaborative Learning Systematic Review^AI Literacy Heptagon 2026

Evidence-synthesis in AI in education

What reviews accomplish

Systematic reviews and meta-analyses in AI in education serve several distinct purposes:

  • Establish the evidence base β€” determining whether AI tools (tutoring, feedback, assessment, chatbots) produce learning gains, and how large those gains are.
  • Map the field and its gaps β€” a scoping review documents what has been studied, where the evidence is concentrated, and where it is missing (e.g., workplace settings, non-English work, failure cases).^AI Vocational Education Training Review
  • Identify moderators and conditions β€” meta-analysis tests whether effects differ by learner population, domain, AI system type, or study design, revealing for whom and under what conditions a tool works.
  • Expose methodological quality β€” reviews routinely find that the field relies on underpowered, pre-experimental, or quasi-experimental designs and immediate post-tests, tempering conclusions.^AI Vocational Education Training Review^Zerkouk Comprehensive Review ITS 2025

Examples from the wiki

AI-era synthesis challenge: productivity vs. learning

Reviews of generative-AI interventions face a distinctive challenge that the wiki's synthesis research highlights: separating productivity gains from durable learning gains. Because generative AI can inflate immediate task performance (homework, assisted practice) without producing learning, meta-analyses must be careful about which outcome they pool. The GenAI-and-programming meta-analysis found large productivity gains but no significant learning gain (g β‰ˆ 0) β€” a clean illustration. Large-scale field studies and unassisted-measure research show that the measured effect depends on whether outcomes are AI-assisted or proctored/unassisted. Reviews should therefore report assisted and unassisted outcomes separately, distinguish performance from learning, and flag studies that measure only immediate AI-supported performance. This connects to AI Ed Evaluation and Summative Assessment.

Strengths and limitations

Strengths:

  • Efficient synthesis of a large, fragmented literature
  • Meta-analysis yields pooled effect estimates, increases statistical power, and detects moderators and heterogeneity
  • Systematic protocols improve transparency and reproducibility over narrative reviews
  • Essential for evidence-based practice and for identifying research gaps

Limitations:

  • Garbage-in/garbage-out β€” the synthesis is only as good as the quality of included studies; weak primary designs yield weak pooled conclusions
  • Publication bias β€” null or negative results are under-published, inflating pooled effects
  • Heterogeneity β€” varied designs, outcome measures, and AI systems make direct pooling hard and can undermine the meaning of a single effect size
  • Rapid obsolescence β€” the AI tool landscape changes quickly, so reviews can date fast
  • Scope constraints β€” single-database or English-only searches may miss relevant work.^AI Collaborative Learning Systematic Review^AI Vocational Education Training Review

Relationship to other methods

Within the wiki's methodological landscape, meta-analysis and systematic review are the synthesis family, complementing primary designs:

  • Primary studies (experiments, surveys, qualitative work, design-based research) generate individual findings; reviews aggregate them. See Research Methods AIED.
  • Effect-size reporting in primary studies (e.g., RCTs) is what makes later meta-analysis possible β€” reviews depend on studies reporting comparable, extractable effect sizes.
  • Evaluation (AI Ed Evaluation, Benchmark) assesses individual systems; reviews assess the literature on systems and interventions.
  • Educational measurement (Educational Measurement, Assessment Validity) concerns the quality of the outcome measures that reviews pool.

Implications for researchers

  1. Report extractable effect sizes. For a literature to be meta-analyzable, primary studies must report comparable effect sizes and adequate methods detail β€” a responsibility of every AIED study.^Research Methods AIED
  2. Follow a transparent protocol. PRISMA-guided search, screening, and appraisal make reviews reproducible and defensible.
  3. Interpret pooled effects cautiously. Attend to heterogeneity, publication bias, and the quality of included studies before drawing strong conclusions.
  4. Use reviews to set the agenda. Reviews' documented gaps (failure cases, workplace settings, non-English and non-indexed work, long-term outcomes) should guide where new primary research is needed.^AI Vocational Education Training Review

Connected Concepts

Connected Articles