On this page

Synthesis: In brief: A World Bank systematic review and meta-analysis pools 191 effect sizes from 14 randomized trials across ten economies to estimate that adaptive and AI-enabled educational technology raises student learning by an average of 0.125 standard deviations relative to traditional instruction — above the median effect for education RCTs and within the range Kraft (2020) calls "large" for field experiments. Crucially, the newer generative-AI tools show no advantage over the adaptive software that preceded them, and gains are driven less by which technology is used than by whether it is embedded in a sound instructional strategy.

The review brings two "generations" of adaptive educational technology into a common framework under common inclusion criteria and on a common effect-size scale: first-generation tools that select from content authored in advance (adaptive computer-assisted learning, intelligent tutoring systems such as Mindspark), and second-generation generative AI tools that generate instructional content at the point of use (Rori in Ghana, GPT-based tutors in Türkiye). Using robust variance estimation (RVE) meta-regression, the authors retain every extracted outcome per study rather than one estimate each — a methodological choice that matters because within-study variation across outcome measures is as wide as variation across studies.

The pooled learning effect of 0.125 sd holds at primary and secondary levels, in high- and middle-income countries, across tutoring systems, computer-assisted learning platforms, and teacher-facing tools, and across both technological generations. The estimated differential for generative over first-generation tools is only 0.022 sd (SE 0.075), an interval wide enough to bound rather than resolve the comparison: the experimental record to date shows no advantage for the newer technology. Gains on socio-emotional outcomes are positive but roughly a quarter the size (0.029 sd, from only five papers), and the review cannot reject the null for teaching practices (three studies). Among AI-powered tutoring interventions specifically, the average effect is 0.12 sd — below the 0.288 sd Nickow et al. (2024) report for human tutoring, but at a small fraction of its cost.

Two features of the evidence base sharply limit what these estimates can justify. The sample is narrow: no included study was conducted in a low-income country, researchers were involved in implementing 16 of 19 interventions, and none was implemented by a government alone — so behavior under public delivery at scale is essentially unobserved. And only two or three studies report per-student costs on a comparable basis, so the field's motivating claim that personalization can be delivered at a fraction of the cost of human tutoring has almost never been measured alongside the effects it is meant to justify.

Key Findings

  • Adaptive and AI-enabled EdTech raises learning by ~0.125 sd on average across 191 effect sizes from 14 RCTs in ten economies — above the median (0.10 sd) for education interventions evaluated by randomized trial, and within the range Kraft (2020) reads as large for broad-achievement field experiments.
  • Generative AI shows no advantage over earlier adaptive software. The differential for second-generation tools is 0.022 sd (SE 0.075), an interval [−0.15, 0.19] wide enough to bound the comparison rather than resolve it; the experimental record to date shows no advantage for the newer technology.
  • Effects are consistent across education levels and income contexts — primary (0.137 sd) and secondary (0.135 sd) are nearly identical, and high- vs. middle-income differences are not statistically significant.
  • AI-powered tutoring averages 0.12 sd, below the 0.288 sd pooled effect for human tutoring (Nickow et al., 2024) but at a small fraction of its cost; the strongest results come from tutoring embedded in a broader instructional strategy with clear objectives, curriculum alignment, and safeguards against misuse.
  • Gains do not extend equally to other outcomes: socio-emotional/behavioral effects are ~0.029 sd (a quarter of the learning effect, five papers), and teaching-practice effects are indistinguishable from zero (three studies).
  • The evidence base is narrow and policy-relevant gaps remain: no low-income-country study, researcher-heavy implementation, no government-alone implementation, and cost data in only ~2–3 of 14 studies.
  • Dosage does not order the estimates — cumulative exposure varies over two orders of magnitude without mapping onto the ranking of effects; within-study measure choice shifts an estimate as much as the study chosen (one ITS's four estimates spanned 0.60 sd).

Connected Concepts

  • Learning Gains — the primary outcome domain and the review's organizing metric
  • Adaptive Learning — the shared feature of all included interventions
  • Intelligent Tutoring — first-generation adaptive tutors and the AI-tutoring subgroup
  • Generative AI — the second-generation tools whose differential is estimated
  • RCT — the inclusion criterion and evidence hierarchy
  • Personalized Learning — the personalization promise motivating the literature
  • Cognitive Offloading — the "effort substitution" mechanism behind the Türkiye harm finding
  • Reducing AI Misuse — the guardrails that removed the harm without improving scores
  • Equity — the digital-divide and no-low-income-setting limitation
  • Digital Divide — infrastructure requirements that bind in the settings with largest deficits
  • K-12 — the preK–12 study population
  • AI Ed Evaluation — meta-analytic evaluation of AI education tools
  • Educational Measurement — standardized effect-size aggregation and RVE methodology

Connected Articles

Citation

Burneo, A., Dinarte-Diaz, L., Lopez, C., & Molina, E. (2026). Can EdTech close learning gaps? Global evidence from digital interventions. World Bank Policy Research Working Paper.

Embed this page

Copy the code below to embed a chromeless version of this page in a learning management system or other website. The embedded view hides the site header, navigation, and footer.