Concept
Transfer of Learning
Transfer of Learning — the extent to which knowledge or skills acquired in one context (e.g., practice with an AI tool) persist and apply in a different context (e.g., independent performance without the tool). In AI in education, transfer is the central open question: whether performance gains students show with AI tools translate into durable learning they can demonstrate without them.
Questions to Consider
- Here's a striking pattern the page documents: students often show immediate gains on AI-assisted tasks, yet those gains can vanish — or even reverse — when the AI is removed. Before reading the explanations, why do you think a tool that clearly helps in the moment could end up leaving students worse off without it?
- Recall something you learned to do with a tutor, calculator, or assistant and then had to do alone. Did the skill carry over, or did you feel dependent on the aid? What was different about the experiences that transferred well versus those that didn't?
- A common intuition is that 'practice is practice' — that doing a task with help builds the same skill as doing it alone. Where might that intuition mislead, especially when the help is an AI that completes the reasoning for you rather than guiding you through it?
- The page draws a distinction between 'effects with' a technology and 'effects of' it — performing better while using the tool versus becoming more capable without it. If you're an instructor, designer, or student, which of these is your real goal, and how would you know you'd achieved it?
- The evidence suggests that how much cognitive work you delegate matters: offloading surface tasks like grammar hurt transfer less than offloading deep reasoning and structure. Think about the last time you used AI on an assignment. Which 'layer' did you delegate, and what does your choice predict about what you'd retain?
- The page proposes conditions that might support positive transfer — pedagogical Guardrails, fading support, calibration to the learner's readiness. If you were designing (or were the user of) an AI learning tool, what would you insist on so that gains while using it become durable ability without it?
Introduction
Transfer of learning is a foundational concern in education research, and AI tools have made it urgent. The defining empirical pattern documented across AI in education studies is a transfer paradox: students using AI typically show immediate, measurable gains on tasks where AI is available, but those gains often fail to persist — or even reverse — when AI is removed and students must demonstrate understanding independently. This pattern implicates Over-Reliance, Cognitive Load Theory, and Metacognition as the mechanisms at work, and connects directly to debates about AI Tutoring design.
The transfer paradox
Students using AI typically show immediate, measurable gains on the tasks where AI is available. Yet when AI is removed:
- Effects become mixed or negative
- Gains often fail to transfer to unassessed settings
- Students may become dependent on the tool at the expense of independent reasoning
The evidence base, synthesized in the Stanford Evidence Base on AI in K-12 review, is consistent across domains:
| Study | Context | Immediate Effect | Transfer Effect | Mechanism |
|---|---|---|---|---|
| Bastani et al. (2025) | High school math | Higher practice grades | ~17% worse on closed-book finals | General-purpose chatbot did the work |
| Chen et al. (2025) | Programming homework | Higher homework scores | No improvement on unassisted exams | Large Language Models (LLMs)-Tutor solved problems for students |
| Lehmann et al. (2025) | Programming | More topics covered | Harmed understanding; widened gaps | General AI for low-prior learners |
| Stadler et al. (2024) | Academic research | Faster task completion | Lower-quality reasoning vs. search | Reduced cognitive engagement |
| Kosmyna et al. (2025) | Essay writing | Higher essay quality | 83% failed to recall their own quotes | Outsourced authorship |
All five studies show a negative or null transfer pattern when general-purpose AI is the intervention.
Mechanisms undermining transfer
Metacognitive displacement. AI completing reasoning reduces opportunities for students to monitor their own understanding and select strategies. Students who used AI were less able to explain their answers when queried. This connects to Metacognition research on self-monitoring and the evidence that structured courses increase metacognitive competence while raw LLM assistants do not.
Germane load suppression. General-purpose AI reduces not just extraneous (distracting) cognitive load but also germane load — the productive mental effort that encodes durable knowledge. Easier practice feels better but stores weaker traces. See Cognitive Load Theory and the distinction between tutoring-specific vs general AI.
Over-reliance / expertise reversal. Novices given answers do not build schemas. General AI provides answers; effective tutoring provides structured guidance. When novices are given expert-level shortcuts, learning is disrupted — the Desirable Difficulties principle in reverse.
Tool-dependent performance. Students may optimize for the specific affordances of the AI tool (prompt engineering, reliance on generated code structure) rather than building domain generalization — a form of cognitive offloading that feels productive but displaces durable learning.
Layer-sensitive offloading and transfer. Chen (2026) directly tests Salomon, Perkins & Globerson's "effects with vs. effects of technology" distinction in GenAI-assisted writing: an eight-week quasi-experiment found open AI collaboration maximized supported-writing performance but produced the lowest independent no-AI near-transfer outcomes, while bounded support with reflection preserved independent competence. Deeper offloading layers (reasoning, structure) predicted worse transfer than surface layers (grammar). This is direct classroom evidence that AI's with-support performance gains do not transfer to of-support independent performance — and that the depth of delegation, not just whether AI is used, shapes transfer.
A complementary, if confounded, instance comes from physics: the Ruhr University Bochum redesign of an introductory nuclear and particle physics course (Mikhasenko et al., 2026) had students successfully complete collaborative, resource-rich research problems with AI assistance, yet those same students averaged 20.6/80 on a conventional unaided written exam, with several serious attempts unable to complete standard calculations. The authors read this as evidence that assisted performance does not automatically transfer to unprompted performance, and their remedy is deliberate design: making the written exam the sole grade determinant, releasing tutorial problems in advance so class time becomes prepared discussion, and adding prerequisite preparation, worked examples and consolidation around the exploratory AI-permitted work.
Conditions supporting positive transfer
The limited evidence suggests transfer is possible when:
-
Pedagogical guardrails are present — step-by-step hints, misconception targeting, Socratic questioning (Bastani et al., 2025 tutoring variant)
-
Traditional strategies are preserved — note-taking paired with AI use improved retention (Kreijkes et al., 2026)
-
AI is used for formative, not summative, practice — scaffolding during learning, not during assessment
-
The practice format is matched to the knowledge being transferred. Rachatasumrit, Koedinger & Carvalho (2025) find that retrieval-practice gains frequently fail to transfer to unfamiliar problems — they strengthen memory for a procedure without enabling its use in new contexts — and that durable generalization to novel applications requires pairing practice with worked examples that support skill induction; the optimal example–problem ratio therefore depends on whether the content is a verbatim fact or a generalizable skill.
-
Learner expertise is calibrated — the tool adapts support to readiness rather than defaulting to full assistance
-
Transfer as the criterion that separates learning from assistance. Yan and Gašević (2026) build their theory of human-AI learning around transfer under reduced support: assisted performance counts as learning only if the capability persists once the support is withdrawn, which makes transfer the test rather than one outcome among several. Their proposition is directional, that requiring source checking or justification during AI-supported work should improve delayed performance, while repeated low-friction delegation without reconstruction should weaken learners' calibration of their own competence.
This aligns with AI Tutoring research showing that tutoring-specific tools with pedagogical guardrails outperform general-purpose chatbots, and with Scaffolding principles about fading support as competence grows.
Unanswered questions
- Time scale: Does transfer improve over weeks/months of use, or does dependence deepen?
- Domain differences: Is transfer better in well-structured domains (math) vs. ill-structured domains (writing)?
- Individual differences: Do high-Prior Knowledge students suffer less transfer loss than novices?
- Skill remediation: Can explicit "AI-off" practice sessions reverse tool dependence?
Connections to related concepts
Transfer of learning connects to Metacognition (self-monitoring of understanding), Cognitive Load Theory (germane vs extraneous load), Desirable Difficulties (productive struggle), Scaffolding (fading support), Over-Reliance (tool dependence), and Sociocultural Learning (general-purpose AI operates outside the ZPD by completing work for students). It is the bridge between assisted performance and genuine learning — the distinction between The Evidence Base on AI in K-12: A 2026 Review and the central question for AI Tutoring effectiveness.
Connected Concepts
- Metacognition
- Desirable Difficulties
- Cognitive Offloading
- Scaffolding
- Sociocultural Learning
- Intelligent Tutoring
- K-12
- Self-Regulated Learning
- Learning Theories
- Productive Failure — Productive Failure
Connected Articles
-
Agentivism: a learning theory for the age of artificial intelligence — A mid-range learning theory for human-AI interaction, with four mechanisms and six testable propositions (Yan and Gašević 2026)
-
Layer-Sensitive Cognitive Offloading in Generative AI-Assisted Writing: Supported Performance and Independent No-AI Outcomes — Layer-sensitive cognitive offloading in GenAI-assisted writing (Chen 2026)
-
Deceptive Overgeneralization: When Adaptive Learning Enables Systematic Misapplication — Deceptive overgeneralization: adaptive mastery can stop practice before learners know when to withhold an action (An, McLaren & Stamper 2026)
-
The critical-thinking paradox in generative AI-integrated learning: distinguishing efficiency from cognitive depth — a differentiated framework and testable propositions — The critical-thinking paradox in GenAI-integrated learning
-
Knowledge without Wisdom: Measuring Misalignment between LLMs and Intended Impact
-
Cognitive offloading and the speedup illusion in human-AI interaction
-
SafeTutors: Benchmarking Pedagogical Safety in AI Tutoring Systems
-
Rethinking Higher Education: From Fixed Curricula to Learnity Graphs
-
Exploring Students' Perceptions of Using Generative AI-Assisted Problem Posing
-
Young People, Learning, and Generative AI: A Rapid Literature Review and Implications for PreK-12 Education — Performance-learning distinction and durable transfer
-
Designing AI systems to support a productive-failure-based learning: insights from adult learners on AI applications — Designing AI Systems to Support Productive-Failure-Based Learning
-
Towards the Pedagogical Steering of Large Language Models for Tutoring: A Case Study with Modeling Productive Failure — Pedagogical Steering of LLMs for Productive Failure
-
AI in Particle Physics Education: Research Problems and Foundational Skills — AI in Particle Physics Education: Research Problems and Foundational Skills
-
From Memorization to Experiential Learning: Reconfiguring Classroom Pedagogy in Management Education through Generative AI — argues that protected classroom simulations can form decision habits that fail outside them
-
Evaluation of pre-trained models for pedagogical assessment of novel AI-assisted educational questions — Evaluation of pre-trained models for pedagogical assessment of novel AI-assisted educational questions