On this page

Learning gains — measurable improvements in student knowledge, skills, or competencies resulting from educational interventions, including AI-assisted instruction. In AI in education research, learning gains serve as the primary outcome measure for evaluating whether AI tools actually improve learning — not just engagement or satisfaction.

Questions to Consider

  • Have you ever felt you learned a lot from an activity, only to fail a test that measured something different? The page distinguishes immediate AI-supported performance from durable learning — how might those two diverge in your own experience?
  • A core finding is that generative AI can inflate scores on AI-assisted homework while lowering scores on proctored, closed-book measures. If you were evaluating whether an AI tool really helps students learn, which outcome would you trust and why?
  • Research shows unguided reliance on AI predicts worse learning gains, while structured use predicts better ones — the same tool, opposite outcomes. What distinguishes 'structured' from 'unguided' use in a real classroom?
  • Hint buttons correlate with reduced learning: more hints, less learning. Have you ever been tempted to reach for a hint or an answer the moment you were stuck? What does that suggest about how much struggle is actually necessary for learning?
  • A large meta-analysis pooled many studies and found AI-enabled EdTech raised learning by a modest amount, with no advantage for generative AI over earlier adaptive tools. How should this cautious, pooled estimate change how you read exciting claims about a single AI product's effectiveness?

Introduction

Learning gains are the ultimate test of any educational technology. In the knowledge base's research, they appear as dependent variables in randomized controlled trials, pre-post comparisons in quasi-experimental studies, and correlational analyses linking AI tool usage to academic outcomes.

Terminology. The outcome sense of achievement — student achievement, academic achievement, learning achievement, prior achievement, achievement gaps — is treated as a synonym for learning gains here, and those phrases link to this page. The felt sense ("a sense of achievement") is a motivational experience rather than a measured outcome; see Motivation and Self-Efficacy. Achievement-goal theory ("achievement goals", goal orientations) is likewise a motivational construct and links to Motivation.

Key findings from the knowledge base:

  • Adaptive pretesting research examines whether GenAI-enabled pretesting produces durable learning gains that persist beyond immediate testing.
  • Meta-analyses of GenAI in programming find positive learning gains from structured AI use but negative effects from unguided reliance — a key distinction between productivity and durable learning.
  • Hint button research shows negative associations between hint abuse and learning gains — more hints correlate with less learning.
  • Instructional guidance studies demonstrate that learning gains depend on HOW AI is used, not just WHETHER it's available.
  • World Bank meta-analysis pools 191 effect sizes from 14 RCTs to estimate that adaptive and AI-enabled EdTech raises learning by ~0.125 sd on average — above the median for education RCTs — while finding no advantage for generative AI over earlier adaptive tools.
  • Gains are content–treatment interactions, not constants. Rachatasumrit et al. (2025) show the optimal example–problem ratio depends on knowledge content: pure retrieval practice yields higher gains for verbatim facts, while example-integrated practice (alternating worked examples and problems) yields higher gains for generalizable skills — direct evidence that "more practice" is not always better and that gains hinge on matching the training schedule to the knowledge component being learned.

The AI-era measurement problem

A central theme in the knowledge base's learning-gains research is that generative AI can inflate apparent performance without producing learning gains — and that the choice of outcome measure determines whether this is visible. Research and large-scale field data show a sharp divergence: AI use improves scores on AI-assisted homework while lowering scores on proctored, closed-book, unassisted measures. Performance-vs-learning research and rapid reviews therefore distinguish immediate AI-supported performance from durable learning, and treat unassisted summative measures (see Summative Assessment) as the reliable signal of genuine learning gains.

What the efficacy research shows

Across the knowledge base's RCTs, meta-analyses, and field studies, a consistent picture of learning efficacy (which AI interventions actually produce learning gains, and how large) emerges:

  • Meta-analytic evidence is broadly positive but conditional. A comprehensive meta-analysis of 53 studies (Dong 2026) finds generative AI generally outperforms traditional approaches on academic achievement, higher-order thinking, and writing — with AI feedback particularly effective — though game-assisted GenAI shows no significant added benefit and gains vary by country and outcome. The GenAI-and-programming meta-analysis finds large productivity gains but no significant learning gain (g ≈ 0), separating task-efficiency from durable learning. A language-learning meta-analysis finds positive but modest learning gains from AI-enhanced embodied robots. For AI literacy specifically, a three-level meta-analysis of 59 studies estimates a large overall effect (g = 0.837) — but the wide prediction interval and the finding that knowledge-focused interventions outperformed those targeting skills, attitudes, or Ethics caution that the outcome measured shapes the apparent gain, echoing the broader point that AI-related gains depend on what and how you assess.
  • Well-designed AI tutors produce real gains. A two-year cluster RCT (Khanmigo) found AI tutoring raised math achievement ~1.3 national percentile ranks per term (~0.06–0.08 SD/school year, ~0.14 SD for a full year), gains resembling practice without AI — demonstrating that engagement, not model capability, is the binding constraint. NUMI showed AI support improved next-attempt correctness after mistakes with more time per question — a "productive slowdown" that builds durable mastery. Virtual tutoring found the binding constraint is take-up and sustained participation, not tutor quality.
  • AI can match human help. ChatGPT-generated help produces learning gains equivalent to human tutor-authored help on mathematics skills — evidence that generative AI can be as efficacious as human Scaffolding when used appropriately.
  • Unguarded AI can harm learning. The guardrail RCT (PNAS 2025) found an unguarded ChatGPT-style tutor raised assisted practice +48% but reduced unassisted exam scores −17%, while a guardrailed (hint-not-answer) tutor eliminated the harm. This is the sharpest demonstration that learning efficacy is design-contingent: the same class of tool can be a strong learning gain or a net harm depending on how it is configured.
  • Perceived vs. actual efficacy diverge. Self-reported performance misaligns with measured performance, and the absent cognitive baseline shows AI-native students overestimate their learning — so efficacy claims based on self-report are unreliable without objective outcome measures. Self-Report Measures collects the cases where reported and measured outcomes come apart.

Takeaway: the weight of evidence supports modest, conditional, and design-dependent learning gains from AI — real when AI is structured to coach rather than answer, guardrailed, and paired with unassisted outcome measures, and absent or negative when it substitutes for the learner's own effort. This is why learning gains as an outcome must be measured with valid, AI-resistant instruments and why AI Ed Evaluation pairs efficacy claims with methodological scrutiny.

  • AI tutoring can match expert human tutoring on standardized-test gains. In a 2,383-participant evaluation, pooled AI tutoring was statistically equivalent to expert human tutoring on GRE learning gains (p = .015), with both above a no-tutoring control; the authors propose cost per percentage point of gain as the unit for comparing tutors (Northcutt et al., 2026).

A field-level map of AI's effects on learning and achievement

The knowledge base's corpus — meta-analyses, RCTs, quasi-experiments, and field studies — supports a nuanced, sometimes contradictory picture. Findings cluster into positive, negative, and conditional effects.

Positive effects (AI can produce genuine learning gains):

Negative effects (AI can reduce or fail to improve learning):

Conditional and mixed effects (context determines direction):

Meta-analytic learning-gain estimates are inflated — read them with caution

The learning-gain numbers that dominate this page's efficacy summary — especially pooled effect sizes from meta-analyses — must be read with a strong caveat: a wave of meta-research shows the field's positive synthesis estimates are inflated by publication bias, construct incoherence, and methodological shortcuts.

  • A study-level meta-meta-analysis of 1,840 effect sizes from 67 meta-analyses finds the publication-bias-adjusted average AI effect is roughly one-third the reported magnitude (SMD ≈ 0.196 vs. a median of 0.67), with a prediction interval spanning large harm to large benefit and no moderator producing consistent gains. Even a single extreme study contributes almost no information given the heterogeneity — meaning more studies of the current type will not settle the question; only high-quality, pre-registered, replication-oriented trials will.
  • A second-order meta-analysis of the synthesis layer itself. Emslander et al. (2026) invert the usual unit of analysis: treating 45 meta-analyses as their data and correcting for study overlap (818 unique primary studies, weighted by uniqueness rather than discarded), they pool 129 meta-analytic effect sizes to g = .57 [.50, .65], with outcome clusters running from .35 for STEM performance to .68 for general academic outcomes. Two results bear on the numbers above. AI type did not moderate the effect (p = .63) — intelligent tutoring systems, chatbots and ChatGPT are statistically indistinguishable at this layer, so the field's pooled gains are not being driven by generative AI as such — and the one clear moderator was educational level (tertiary .61 vs. .30 for younger samples), which is partly a study-design artifact given that 74% of the corpus tested tertiary students and K-12 evidence is comparatively thin. Its quality appraisal is the sharpest in the corpus: the 45 meta-analyses averaged 9.3 of 17 on an AMSTAR-adapted scale, none preregistered, and only 23 of 45 had 80% power to detect the effect they themselves reported. Its two publication-bias tests also disagree — a symmetric funnel plot (Kendall's τ = −.01, p = .868) against PET-PEESE (B = 0.9, p = .009) indicating inflation — so the SOMA's own medium effect arrives with an unresolved bias cloud over it.
  • A forensic audit of 14 high-impact AIED meta-analyses finds none provided a valid basis for its pooled learning-gain claim: none had a coherent outcome construct, all had unresolved extreme heterogeneity (I² ranged from 77.2% to 94.4% across the 13 meta-analyses that reported it, and 12 of those 13 exceeded 80%), twelve treated dependent effect sizes as independent, and none validly assessed publication bias. A majority of randomly vetted primary studies were mismatched to the meta-analytic claim. The errors reached policy: two of the audited meta-analyses treated study sample size as class size and, on the strength of a miscalculated primary study, concluded that 21 to 40 students is the ideal intervention size and recommended that GenAI interventions be designed for that number.
  • A 2026 STEM synthesis that fixes two of the audited defects — and still illustrates a third. Doğan and colleagues (2026) pooled 35 experimental and quasi-experimental STEM studies with one effect size per study to avoid the dependent-effect-size problem, and ran a full publication-bias battery (funnel plot, Begg's and Egger's tests, trim-and-fill, Rosenthal's fail-safe N = 2404) rather than asserting symmetry — the two shortcuts the audit above found in every meta-analysis it examined. Two cautions survive. First, the authors applied no formal quality appraisal and treated the inclusion criteria as the rigor threshold, so studies of very different designs entered the pool unweighted by quality. Second, the headline heterogeneity depends entirely on which model you read: the same 35 studies are reported as I² = 82.98% under the fixed-effect model and I² = 15.75% under the random-effects model, so a reader who encounters the numbers without the model label can conclude the corpus is homogeneous when it is not. Its pooled estimate (g = 0.670, 95% CI [0.491, 0.848]) is therefore the kind of figure the section above describes: better built than most, and still an upper bound.
  • A media/methods critique shows many "learning gains" were not measured validly — outcomes were often self-reported skills or performance measured during AI assistance rather than durable, unassisted learning.

Bottom line for the gains numbers above: treat large pooled AI effect sizes as upper bounds, not point estimates. Prefer the learning-gain evidence from well-designed RCTs and field studies with unassisted, standardized outcome measures (the guardrail RCT, the Khanmigo and NUMI experiments, and the World Bank EdTech meta-analysis cited above), and read meta-analytic gains as provisional and likely over-stated until synthesis quality improves. The problem also outlives correction: of the papers published after one audited meta-analysis was retracted, 60% still cited it as authoritative support for large ChatGPT learning gains and none acknowledged the retraction, so a withdrawn estimate keeps propagating. This is why the Limitations in AIEd Research page now documents the meta-analytic evidence crisis in detail.

Measuring what matters

Learning gains connect to Assessment Validity — if assessments fail to capture deeper understanding, learning gain measures are misleading. They also intersect with Over-Reliance and Cognitive Offloading, where apparent performance improvements may mask learning losses, and with RCT (randomized trials as the gold-standard design for detecting causal learning gains), and with Meta-Analysis and Systematic Review (pooling effect sizes across studies to establish the field's efficacy evidence).

  • Significant pre/post gains from mistake-based AI Pedagogies and Teaching Strategies: Hosseini (2026)'s database design course (n=13) showed large, significant learning gains on identical pre/post items (mean 4.25→6.83/7, Cohen's d=1.49, p<.001), with gains uncorrelated with prior AI or database confidence — the AI-integrated critique-refinement design benefited students regardless of initial perceptions.

Connected Concepts

Connected Articles

Connected FAQs

Embed this page

Copy the code below to embed a chromeless version of this page in a learning management system or other website. The embedded view hides the site header, navigation, and footer.