Concept
Learning Gains
Learning gains — measurable improvements in student knowledge, skills, or competencies resulting from educational interventions, including AI-assisted instruction. In AI in education research, learning gains serve as the primary outcome measure for evaluating whether AI tools actually improve learning — not just engagement or satisfaction.
Questions to Consider
- Have you ever felt you learned a lot from an activity, only to fail a test that measured something different? The page distinguishes immediate AI-supported performance from durable learning — how might those two diverge in your own experience?
- A core finding is that generative AI can inflate scores on AI-assisted homework while lowering scores on proctored, closed-book measures. If you were evaluating whether an AI tool really helps students learn, which outcome would you trust and why?
- Research shows unguided reliance on AI predicts worse learning gains, while structured use predicts better ones — the same tool, opposite outcomes. What distinguishes 'structured' from 'unguided' use in a real classroom?
- Hint buttons correlate with reduced learning: more hints, less learning. Have you ever been tempted to reach for a hint or an answer the moment you were stuck? What does that suggest about how much struggle is actually necessary for learning?
- A large meta-analysis pooled many studies and found AI-enabled EdTech raised learning by a modest amount, with no advantage for generative AI over earlier adaptive tools. How should this cautious, pooled estimate change how you read exciting claims about a single AI product's effectiveness?
Introduction
Learning gains are the ultimate test of any educational technology. In the knowledge base's research, they appear as dependent variables in randomized controlled trials, pre-post comparisons in quasi-experimental studies, and correlational analyses linking AI tool usage to academic outcomes.
Terminology. The outcome sense of achievement — student achievement, academic achievement, learning achievement, prior achievement, achievement gaps — is treated as a synonym for learning gains here, and those phrases link to this page. The felt sense ("a sense of achievement") is a motivational experience rather than a measured outcome; see Motivation and Self-Efficacy. Achievement-goal theory ("achievement goals", goal orientations) is likewise a motivational construct and links to Motivation.
Key findings from the knowledge base:
- Adaptive pretesting research examines whether GenAI-enabled pretesting produces durable learning gains that persist beyond immediate testing.
- Meta-analyses of GenAI in programming find positive learning gains from structured AI use but negative effects from unguided reliance — a key distinction between productivity and durable learning.
- Hint button research shows negative associations between hint abuse and learning gains — more hints correlate with less learning.
- Instructional guidance studies demonstrate that learning gains depend on HOW AI is used, not just WHETHER it's available.
- World Bank meta-analysis pools 191 effect sizes from 14 RCTs to estimate that adaptive and AI-enabled EdTech raises learning by ~0.125 sd on average — above the median for education RCTs — while finding no advantage for generative AI over earlier adaptive tools.
- Gains are content–treatment interactions, not constants. Rachatasumrit et al. (2025) show the optimal example–problem ratio depends on knowledge content: pure retrieval practice yields higher gains for verbatim facts, while example-integrated practice (alternating worked examples and problems) yields higher gains for generalizable skills — direct evidence that "more practice" is not always better and that gains hinge on matching the training schedule to the knowledge component being learned.
The AI-era measurement problem
A central theme in the knowledge base's learning-gains research is that generative AI can inflate apparent performance without producing learning gains — and that the choice of outcome measure determines whether this is visible. Research and large-scale field data show a sharp divergence: AI use improves scores on AI-assisted homework while lowering scores on proctored, closed-book, unassisted measures. Performance-vs-learning research and rapid reviews therefore distinguish immediate AI-supported performance from durable learning, and treat unassisted summative measures (see Summative Assessment) as the reliable signal of genuine learning gains.
What the efficacy research shows
Across the knowledge base's RCTs, meta-analyses, and field studies, a consistent picture of learning efficacy (which AI interventions actually produce learning gains, and how large) emerges:
- Meta-analytic evidence is broadly positive but conditional. A comprehensive meta-analysis of 53 studies (Dong 2026) finds generative AI generally outperforms traditional approaches on academic achievement, higher-order thinking, and writing — with AI feedback particularly effective — though game-assisted GenAI shows no significant added benefit and gains vary by country and outcome. The GenAI-and-programming meta-analysis finds large productivity gains but no significant learning gain (g ≈ 0), separating task-efficiency from durable learning. A language-learning meta-analysis finds positive but modest learning gains from AI-enhanced embodied robots. For AI literacy specifically, a three-level meta-analysis of 59 studies estimates a large overall effect (g = 0.837) — but the wide prediction interval and the finding that knowledge-focused interventions outperformed those targeting skills, attitudes, or Ethics caution that the outcome measured shapes the apparent gain, echoing the broader point that AI-related gains depend on what and how you assess.
- Well-designed AI tutors produce real gains. A two-year cluster RCT (Khanmigo) found AI tutoring raised math achievement ~1.3 national percentile ranks per term (~0.06–0.08 SD/school year, ~0.14 SD for a full year), gains resembling practice without AI — demonstrating that engagement, not model capability, is the binding constraint. NUMI showed AI support improved next-attempt correctness after mistakes with more time per question — a "productive slowdown" that builds durable mastery. Virtual tutoring found the binding constraint is take-up and sustained participation, not tutor quality.
- AI can match human help. ChatGPT-generated help produces learning gains equivalent to human tutor-authored help on mathematics skills — evidence that generative AI can be as efficacious as human Scaffolding when used appropriately.
- Unguarded AI can harm learning. The guardrail RCT (PNAS 2025) found an unguarded ChatGPT-style tutor raised assisted practice +48% but reduced unassisted exam scores −17%, while a guardrailed (hint-not-answer) tutor eliminated the harm. This is the sharpest demonstration that learning efficacy is design-contingent: the same class of tool can be a strong learning gain or a net harm depending on how it is configured.
- Perceived vs. actual efficacy diverge. Self-reported performance misaligns with measured performance, and the absent cognitive baseline shows AI-native students overestimate their learning — so efficacy claims based on self-report are unreliable without objective outcome measures. Self-Report Measures collects the cases where reported and measured outcomes come apart.
Takeaway: the weight of evidence supports modest, conditional, and design-dependent learning gains from AI — real when AI is structured to coach rather than answer, guardrailed, and paired with unassisted outcome measures, and absent or negative when it substitutes for the learner's own effort. This is why learning gains as an outcome must be measured with valid, AI-resistant instruments and why AI Ed Evaluation pairs efficacy claims with methodological scrutiny.
- AI tutoring can match expert human tutoring on standardized-test gains. In a 2,383-participant evaluation, pooled AI tutoring was statistically equivalent to expert human tutoring on GRE learning gains (p = .015), with both above a no-tutoring control; the authors propose cost per percentage point of gain as the unit for comparing tutors (Northcutt et al., 2026).
A field-level map of AI's effects on learning and achievement
The knowledge base's corpus — meta-analyses, RCTs, quasi-experiments, and field studies — supports a nuanced, sometimes contradictory picture. Findings cluster into positive, negative, and conditional effects.
Positive effects (AI can produce genuine learning gains):
- Meta-analytic evidence is broadly positive but conditional. A 53-study meta-analysis (Dong 2026) finds generative AI generally outperforms traditional approaches on academic achievement, higher-order thinking, and writing, with AI Feedback especially effective. A meta-analysis of 29 experiments (Does Generative Artificial Intelligence Improve Students' Higher-Order Thinking? A Meta-Analysis Based on 29 Experiments and Quasi-Experiments) finds a moderate positive effect on higher-order thinking, strongest for Problem Solving, but limited for Creativity.
- Tutoring-specific AI reliably outperforms general-purpose AI. The Stanford SCALE review and the umbrella review (Generative AI in Higher Education: A Systematic Review of Opportunities, Challenges, and Pedagogical Innovations (2022–2025)) converge: pedagogically designed agents with hints and step-by-step scaffolding produce real gains where open chatbots often do not.
- Well-designed AI tutors produce real, measurable gains. The two-year Khanmigo RCT (One Click Away: AI Tutoring with Khanmigo in a Two-Year School Experiment) raised math achievement ~1.3 national percentile ranks per term; NUMI improved next-attempt correctness after mistakes; virtual tutoring showed take-up is the binding constraint.
- AI can match human help. ChatGPT-generated help produces learning gains equivalent to human tutor-authored help on math skills.
- Feedback-focused AI works. AI feedback is among the most effective GenAI applications; AI feedback for writing improves critical thinking when coupled with instruction. In a field experiment, Geschwind et al. (2026) found students receiving individual GPT-4 feedback on open-ended tasks showed the largest content learning gains (~0.11, p < 0.10; rising to 0.16 among those who actually received prior feedback) — an effect driven by reliable, consistent AI provision rather than inherent superiority, since peer outcomes matched AI when high-quality peer feedback was actually delivered. LLM critique partners in writing extend this to iterative, collaborative feedback: Oppenheimer, Cash & Connell Pensky (2025) found significant gains across a semester on argumentative writing, prompt engineering, and response-to-AI-feedback quality (all p < .001, roughly a full standard deviation per dimension), with gains appearing even on essays written without LLM support — evidence of durable skill rather than mere tool dependency, though the lack of a control condition limits causal attribution.
- Structured scaffolds yield gains. Scaffolded self-regulated feedback and interaction design show small but significant gains when AI is designed to coach.
- AI in collaborative and problem-based contexts helps when structured. AI-enhanced PBL, AI-assisted collaborative learning, and AI-designed cooperative techniques report meaningful gains.
Negative effects (AI can reduce or fail to improve learning):
- Unguarded AI can harm learning. The PNAS guardrail RCT (Generative AI without guardrails can harm learning: Evidence from high school mathematics) found an unguarded ChatGPT-style tutor raised assisted practice +48% but reduced unassisted exam scores −17%; Guardrails eliminated the harm.
- Faster completion ≠ learning. GenAI reduced study time on math problems and the knowledge they build; performance-vs-learning research shows AI inflates AI-assisted performance while lowering proctored, closed-book, unassisted scores.
- Homework outsourcing harms learning. Field data show a generative-AI learning penalty when students outsource homework.
- Large Language Models (LLMs) reliance correlates with lower grades. Jošt et al. found significant negative correlations between LLM use for code generation (rho = −0.305) and debugging (rho = −0.360) and final grades.
- AI can harm teaching and achievement for some. A teacher-facing GenAI RCT found below-median teachers' students lost ground (−0.129 SD).
- Hint abuse correlates with less learning. Hint button research shows more hints, used unproductively, correlate with less learning.
- Meta-analytic programming gains are illusory. The GenAI-and-programming meta-analysis finds large productivity gains but no significant learning gain (g ≈ 0), separating task-efficiency from durable learning.
Conditional and mixed effects (context determines direction):
- Design is decisive. The same tool class can be a strong gain or a net harm depending on configuration (Generative AI without guardrails can harm learning: Evidence from high school mathematics, The Evidence Base on AI in K-12: A 2026 Review).
- How AI is used matters more than whether. Use for explanations was benign; use for code generation was harmful; over-reliance erodes gains.
- Duration and self-AI Regulation in Education moderate effects. Effects were strongest at 8–16 weeks and for learners with higher Self-Regulated Learning.
- AI-IBL supports creativity but not necessarily problem-solving. Mujib et al. improved creative performance and attitudes but not critical problem-solving.
- Student experience diverges from measured gains. AI-native students overestimate their learning, and self-report misaligns with performance — perceptions of gains are not reliable evidence of them.
Meta-analytic learning-gain estimates are inflated — read them with caution
The learning-gain numbers that dominate this page's efficacy summary — especially pooled effect sizes from meta-analyses — must be read with a strong caveat: a wave of meta-research shows the field's positive synthesis estimates are inflated by publication bias, construct incoherence, and methodological shortcuts.
- A study-level meta-meta-analysis of 1,840 effect sizes from 67 meta-analyses finds the publication-bias-adjusted average AI effect is roughly one-third the reported magnitude (SMD ≈ 0.196 vs. a median of 0.67), with a prediction interval spanning large harm to large benefit and no moderator producing consistent gains. Even a single extreme study contributes almost no information given the heterogeneity — meaning more studies of the current type will not settle the question; only high-quality, pre-registered, replication-oriented trials will.
- A second-order meta-analysis of the synthesis layer itself. Emslander et al. (2026) invert the usual unit of analysis: treating 45 meta-analyses as their data and correcting for study overlap (818 unique primary studies, weighted by uniqueness rather than discarded), they pool 129 meta-analytic effect sizes to g = .57 [.50, .65], with outcome clusters running from .35 for STEM performance to .68 for general academic outcomes. Two results bear on the numbers above. AI type did not moderate the effect (p = .63) — intelligent tutoring systems, chatbots and ChatGPT are statistically indistinguishable at this layer, so the field's pooled gains are not being driven by generative AI as such — and the one clear moderator was educational level (tertiary .61 vs. .30 for younger samples), which is partly a study-design artifact given that 74% of the corpus tested tertiary students and K-12 evidence is comparatively thin. Its quality appraisal is the sharpest in the corpus: the 45 meta-analyses averaged 9.3 of 17 on an AMSTAR-adapted scale, none preregistered, and only 23 of 45 had 80% power to detect the effect they themselves reported. Its two publication-bias tests also disagree — a symmetric funnel plot (Kendall's τ = −.01, p = .868) against PET-PEESE (B = 0.9, p = .009) indicating inflation — so the SOMA's own medium effect arrives with an unresolved bias cloud over it.
- A forensic audit of 14 high-impact AIED meta-analyses finds none provided a valid basis for its pooled learning-gain claim: none had a coherent outcome construct, all had unresolved extreme heterogeneity (I² ranged from 77.2% to 94.4% across the 13 meta-analyses that reported it, and 12 of those 13 exceeded 80%), twelve treated dependent effect sizes as independent, and none validly assessed publication bias. A majority of randomly vetted primary studies were mismatched to the meta-analytic claim. The errors reached policy: two of the audited meta-analyses treated study sample size as class size and, on the strength of a miscalculated primary study, concluded that 21 to 40 students is the ideal intervention size and recommended that GenAI interventions be designed for that number.
- A 2026 STEM synthesis that fixes two of the audited defects — and still illustrates a third. Doğan and colleagues (2026) pooled 35 experimental and quasi-experimental STEM studies with one effect size per study to avoid the dependent-effect-size problem, and ran a full publication-bias battery (funnel plot, Begg's and Egger's tests, trim-and-fill, Rosenthal's fail-safe N = 2404) rather than asserting symmetry — the two shortcuts the audit above found in every meta-analysis it examined. Two cautions survive. First, the authors applied no formal quality appraisal and treated the inclusion criteria as the rigor threshold, so studies of very different designs entered the pool unweighted by quality. Second, the headline heterogeneity depends entirely on which model you read: the same 35 studies are reported as I² = 82.98% under the fixed-effect model and I² = 15.75% under the random-effects model, so a reader who encounters the numbers without the model label can conclude the corpus is homogeneous when it is not. Its pooled estimate (g = 0.670, 95% CI [0.491, 0.848]) is therefore the kind of figure the section above describes: better built than most, and still an upper bound.
- A media/methods critique shows many "learning gains" were not measured validly — outcomes were often self-reported skills or performance measured during AI assistance rather than durable, unassisted learning.
Bottom line for the gains numbers above: treat large pooled AI effect sizes as upper bounds, not point estimates. Prefer the learning-gain evidence from well-designed RCTs and field studies with unassisted, standardized outcome measures (the guardrail RCT, the Khanmigo and NUMI experiments, and the World Bank EdTech meta-analysis cited above), and read meta-analytic gains as provisional and likely over-stated until synthesis quality improves. The problem also outlives correction: of the papers published after one audited meta-analysis was retracted, 60% still cited it as authoritative support for large ChatGPT learning gains and none acknowledged the retraction, so a withdrawn estimate keeps propagating. This is why the Limitations in AIEd Research page now documents the meta-analytic evidence crisis in detail.
Measuring what matters
Learning gains connect to Assessment Validity — if assessments fail to capture deeper understanding, learning gain measures are misleading. They also intersect with Over-Reliance and Cognitive Offloading, where apparent performance improvements may mask learning losses, and with RCT (randomized trials as the gold-standard design for detecting causal learning gains), and with Meta-Analysis and Systematic Review (pooling effect sizes across studies to establish the field's efficacy evidence).
- Significant pre/post gains from mistake-based AI Pedagogies and Teaching Strategies: Hosseini (2026)'s database design course (n=13) showed large, significant learning gains on identical pre/post items (mean 4.25→6.83/7, Cohen's d=1.49, p<.001), with gains uncorrelated with prior AI or database confidence — the AI-integrated critique-refinement design benefited students regardless of initial perceptions.
Connected Concepts
- RCT
- Meta-Analysis and Systematic Review
- Formative Assessment
- Summative Assessment
- Cognitive Offloading
- Math Education
- Human-in-the-Loop
- Affective Tutoring
- Theory Development in AI in Education — Theory Development in AI in Education
- Self-Report Measures
- Research Methods in AIED — Research Methods in AIED (DBR section)
- Social-Emotional Learning — Social-Emotional Learning
Connected Articles
- Distinguishing performance gains from learning when using generative AI — why assisted performance is not a learning outcome (Yan et al. 2025)
- AI Tutoring is Not a Monolith: What We Actually Know — AI Tutoring is Not a Monolith: What We Actually Know (Stanford SCALE/NSSA brief)
- Virtual Tutoring with Computer-Assisted Learning: An Experiment in Take-Up and Learning — Virtual tutoring with CAL: an experiment in take-up and learning
- Making AI Tutoring Productive: Evidence from a Mastery-Based Math Practice Experiment — Making AI tutoring productive: mastery-based math practice
- One Click Away: AI Tutoring with Khanmigo in a Two-Year School Experiment — One Click Away: Khanmigo in a two-year school experiment
- The StudyChat Dataset: Analyzing Student Dialogues With ChatGPT in an Artificial Intelligence Course — The StudyChat dataset of student–LLM dialogues in an AI course
- How to Assess AI Literacy: Misalignment Between Self-Reported and Objective-Based Measures — AI Literacy Assessment: Self-Reported vs Performance Misalignment
- Faster Completion, Less Learning: Generative AI Reduced Study Time on Math Problems and the Knowledge They Build — Faster Completion, Less Learning: Generative AI Reduced Study Time on Math Problems and the Knowledge They Build
- A meta-analysis of the effect of generative AI on productivity and learning in programming — A meta-analysis of the effect of generative AI on productivity and learning in programming
- AI Literacy Interventions in Education: A Meta-Analysis of Effects and Moderators — Meta-analysis of AI literacy intervention effects
- ChatGPT-generated help produces learning gains equivalent to human tutor-authored help on mathematics skills — ChatGPT help produces learning gains equivalent to human tutor help
- Generative AI without guardrails can harm learning: Evidence from high school mathematics — Generative AI without guardrails can harm learning (PNAS 2025 RCT)
- The Absent Cognitive Baseline: Theorizing a Structural Gap in AI-Native College Students' Academic Self-Assessment — The Absent Cognitive Baseline: Theorizing a Structural Gap in AI-Native College Students' Academic Self-Assessment
- OECD Digital Education Outlook 2026: Exploring Effective Uses of Generative AI in Education — OECD Digital Education Outlook 2026
- How AI Is Changing Teaching Workflows — How AI Is Changing Teaching Workflows
- Multimodality and Social Interactions in AI-Enhanced Embodied Robot-Assisted Language Learning: A Meta-Analysis — Meta-analysis of AI-enhanced embodied robot-assisted language learning
- Generative AI technologies and educational outcomes: a comprehensive meta-analysis comparing traditional and AI-driven approaches
- Young People, Learning, and Generative AI: A Rapid Literature Review and Implications for PreK-12 Education — Immediate performance vs durable learning distinction
- The Generative AI Learning Penalty: Evidence from Chinese Secondary Education — The generative AI learning penalty: homework outsourcing harms learning
- Artificial Intelligence Agents in Computer-Supported Collaborative Learning: A Systematic Literature Review — AI agents in computer-supported collaborative learning review
- Design-Based Research for Developing an AI-Assisted Collaborative Learning Model to Enhance Critical Thinking and Problem-Solving Skills in Higher Education — AI-Assisted Collaborative Learning model DBR (critical thinking +24.1%, problem-solving gains)
- The Evidence Base on AI in K-12: A 2026 Review — Tutoring-specific vs general AI
- The Impact of Large Language Models on Programming Education and Student Learning Outcomes — LLM reliance and grades in coding (negative correlations)
- Generative AI Can Harm Teaching — Generative AI can harm teaching (RCT)
- Does Generative Artificial Intelligence Improve Students' Higher-Order Thinking? A Meta-Analysis Based on 29 Experiments and Quasi-Experiments — GenAI and higher-order thinking meta-analysis
- Evaluating the Impact of AI-Supported Inquiry-Based Learning on Students' Creative Mathematical Performance, Critical Problem-Solving Skills, and Attitudes Toward Mathematics — AI-supported IBL and creative math performance
- AI-Enhanced Problem-Based Learning Framework: Integrating ChatGPT as Adaptive Scaffolding to Improve Critical Thinking and Personalized Learning — AI-enhanced PBL scaffolding gains
- Artificial intelligence assisted design of a novel cooperative learning technique for higher education — AI-designed cooperative learning technique
- Fostering feedback literacy by scaffolding self-regulated feedback: a comparative study of GenAI and human peers — Scaffolded self-regulated feedback gains
- Patterns of Learner-AI Interaction and Academic Performance in an Object-Oriented Programming Course — Interaction patterns and learning gains in OOP
- Using AI-Generated Feedback to Improve Critical Thinking and Writing Proficiency — AI feedback and critical thinking in writing
- The Pedagogy of AI Mistakes: Fostering Higher-Order Thinking — The Pedagogy of AI Mistakes: Fostering Higher-Order Thinking (Hosseini 2026)
- Intelligent tutoring in dynamic domains: a graph-based system for comparative analysis of adaptive algorithms — Graph-Based Intelligent Tutoring for Dynamic Domains (2026)
- Exploring the effect of computational thinking levels on students' learning performance, cognition, and behavior when — Computational Thinking Levels and AI Coding Assistants (2026)
- Adaptive Scaffolding for Cognitive Engagement in an Intelligent Tutoring System — Adaptive ICAP scaffolding in an ITS (BKT vs DRL)
- Can EdTech Close Learning Gaps? Global Evidence from Digital Interventions — Meta-analysis: adaptive/AI EdTech raises learning ~0.125 sd
- GPT-4 feedback increases student activation and learning outcomes in higher education
- Evidence and Theory for why the Best Example-Problem Ratio To Optimize Learning Gain Depends on Knowledge Content
- You've Got AI Friend in Me: LLMs as Collaborative Learning Partners
- ChatGPT in Education: An Effect in Search of a Cause — ChatGPT in Education: An Effect in Search of a Cause (media-comparison critique of gains measures)
- Effect of Artificial Intelligence on Learning: A Meta-Meta-Analysis — Meta-meta-analysis: bias-adjusted AI learning-gain effects ~1/3 of reported size
- Presumed Effective: The Manufacturing of an Evidence Base for AI-in-Education Through Flawed Meta-Analysis — Presumed Effective: audit of flawed AIED meta-analyses
- The Impact of Artificial Intelligence-Supported Instruction on Student Learning in STEM: A Systematic Review and Meta-Analysis — A STEM synthesis that controlled dependent effect sizes and tested publication bias, but applied no quality appraisal (Doğan et al. 2026)
- What Do We Know About the Effects of Artificial Intelligence in Education? A Second-Order Meta-Analysis — Second-order meta-analysis of 45 AI-in-education meta-analyses, with overlap-corrected effects and a quality appraisal of the synthesis layer (Emslander et al. 2026)
- StudentBench: AI and human tutoring yield equivalent GRE learning gains — StudentBench: AI and human tutoring yield equivalent GRE learning gains
Connected FAQs
- What Are the Top 10 Findings from AI in Education Research That Instructors Should Know About?
- What Are Notable Gaps in the Research Literature on AI in Education?
- Does Using AI Actually Help My Students Learn?
- What Measures and Research Methods Can an Instructor Use to Evaluate AI-Related Interventions?