On this page

Transfer of Learning — the extent to which knowledge or skills acquired in one context (e.g., practice with an AI tool) persist and apply in a different context (e.g., independent performance without the tool). In AI in education, transfer is the central open question: whether performance gains students show with AI tools translate into durable learning they can demonstrate without them.

Questions to Consider

  • Here's a striking pattern the page documents: students often show immediate gains on AI-assisted tasks, yet those gains can vanish — or even reverse — when the AI is removed. Before reading the explanations, why do you think a tool that clearly helps in the moment could end up leaving students worse off without it?
  • Recall something you learned to do with a tutor, calculator, or assistant and then had to do alone. Did the skill carry over, or did you feel dependent on the aid? What was different about the experiences that transferred well versus those that didn't?
  • A common intuition is that 'practice is practice' — that doing a task with help builds the same skill as doing it alone. Where might that intuition mislead, especially when the help is an AI that completes the reasoning for you rather than guiding you through it?
  • The page draws a distinction between 'effects with' a technology and 'effects of' it — performing better while using the tool versus becoming more capable without it. If you're an instructor, designer, or student, which of these is your real goal, and how would you know you'd achieved it?
  • The evidence suggests that how much cognitive work you delegate matters: offloading surface tasks like grammar hurt transfer less than offloading deep reasoning and structure. Think about the last time you used AI on an assignment. Which 'layer' did you delegate, and what does your choice predict about what you'd retain?
  • The page proposes conditions that might support positive transfer — pedagogical Guardrails, fading support, calibration to the learner's readiness. If you were designing (or were the user of) an AI learning tool, what would you insist on so that gains while using it become durable ability without it?

Introduction

Transfer of learning is a foundational concern in education research, and AI tools have made it urgent. The defining empirical pattern documented across AI in education studies is a transfer paradox: students using AI typically show immediate, measurable gains on tasks where AI is available, but those gains often fail to persist — or even reverse — when AI is removed and students must demonstrate understanding independently. This pattern implicates Over-Reliance, Cognitive Load Theory, and Metacognition as the mechanisms at work, and connects directly to debates about AI Tutoring design.

The transfer paradox

Students using AI typically show immediate, measurable gains on the tasks where AI is available. Yet when AI is removed:

  • Effects become mixed or negative
  • Gains often fail to transfer to unassessed settings
  • Students may become dependent on the tool at the expense of independent reasoning

The evidence base, synthesized in the Stanford Evidence Base on AI in K-12 review, is consistent across domains:

Study Context Immediate Effect Transfer Effect Mechanism
Bastani et al. (2025) High school math Higher practice grades ~17% worse on closed-book finals General-purpose chatbot did the work
Chen et al. (2025) Programming homework Higher homework scores No improvement on unassisted exams Large Language Models (LLMs)-Tutor solved problems for students
Lehmann et al. (2025) Programming More topics covered Harmed understanding; widened gaps General AI for low-prior learners
Stadler et al. (2024) Academic research Faster task completion Lower-quality reasoning vs. search Reduced cognitive engagement
Kosmyna et al. (2025) Essay writing Higher essay quality 83% failed to recall their own quotes Outsourced authorship

All five studies show a negative or null transfer pattern when general-purpose AI is the intervention.

Mechanisms undermining transfer

Metacognitive displacement. AI completing reasoning reduces opportunities for students to monitor their own understanding and select strategies. Students who used AI were less able to explain their answers when queried. This connects to Metacognition research on self-monitoring and the evidence that structured courses increase metacognitive competence while raw LLM assistants do not.

Germane load suppression. General-purpose AI reduces not just extraneous (distracting) cognitive load but also germane load — the productive mental effort that encodes durable knowledge. Easier practice feels better but stores weaker traces. See Cognitive Load Theory and the distinction between tutoring-specific vs general AI.

Over-reliance / expertise reversal. Novices given answers do not build schemas. General AI provides answers; effective tutoring provides structured guidance. When novices are given expert-level shortcuts, learning is disrupted — the Desirable Difficulties principle in reverse.

Tool-dependent performance. Students may optimize for the specific affordances of the AI tool (prompt engineering, reliance on generated code structure) rather than building domain generalization — a form of cognitive offloading that feels productive but displaces durable learning.

Layer-sensitive offloading and transfer. Chen (2026) directly tests Salomon, Perkins & Globerson's "effects with vs. effects of technology" distinction in GenAI-assisted writing: an eight-week quasi-experiment found open AI collaboration maximized supported-writing performance but produced the lowest independent no-AI near-transfer outcomes, while bounded support with reflection preserved independent competence. Deeper offloading layers (reasoning, structure) predicted worse transfer than surface layers (grammar). This is direct classroom evidence that AI's with-support performance gains do not transfer to of-support independent performance — and that the depth of delegation, not just whether AI is used, shapes transfer.

A complementary, if confounded, instance comes from physics: the Ruhr University Bochum redesign of an introductory nuclear and particle physics course (Mikhasenko et al., 2026) had students successfully complete collaborative, resource-rich research problems with AI assistance, yet those same students averaged 20.6/80 on a conventional unaided written exam, with several serious attempts unable to complete standard calculations. The authors read this as evidence that assisted performance does not automatically transfer to unprompted performance, and their remedy is deliberate design: making the written exam the sole grade determinant, releasing tutorial problems in advance so class time becomes prepared discussion, and adding prerequisite preparation, worked examples and consolidation around the exploratory AI-permitted work.

Conditions supporting positive transfer

The limited evidence suggests transfer is possible when:

  • Pedagogical guardrails are present — step-by-step hints, misconception targeting, Socratic questioning (Bastani et al., 2025 tutoring variant)

  • Traditional strategies are preserved — note-taking paired with AI use improved retention (Kreijkes et al., 2026)

  • AI is used for formative, not summative, practice — scaffolding during learning, not during assessment

  • The practice format is matched to the knowledge being transferred. Rachatasumrit, Koedinger & Carvalho (2025) find that retrieval-practice gains frequently fail to transfer to unfamiliar problems — they strengthen memory for a procedure without enabling its use in new contexts — and that durable generalization to novel applications requires pairing practice with worked examples that support skill induction; the optimal example–problem ratio therefore depends on whether the content is a verbatim fact or a generalizable skill.

  • Learner expertise is calibrated — the tool adapts support to readiness rather than defaulting to full assistance

  • Transfer as the criterion that separates learning from assistance. Yan and Gašević (2026) build their theory of human-AI learning around transfer under reduced support: assisted performance counts as learning only if the capability persists once the support is withdrawn, which makes transfer the test rather than one outcome among several. Their proposition is directional, that requiring source checking or justification during AI-supported work should improve delayed performance, while repeated low-friction delegation without reconstruction should weaken learners' calibration of their own competence.

This aligns with AI Tutoring research showing that tutoring-specific tools with pedagogical guardrails outperform general-purpose chatbots, and with Scaffolding principles about fading support as competence grows.

Unanswered questions

  1. Time scale: Does transfer improve over weeks/months of use, or does dependence deepen?
  2. Domain differences: Is transfer better in well-structured domains (math) vs. ill-structured domains (writing)?
  3. Individual differences: Do high-Prior Knowledge students suffer less transfer loss than novices?
  4. Skill remediation: Can explicit "AI-off" practice sessions reverse tool dependence?

Transfer of learning connects to Metacognition (self-monitoring of understanding), Cognitive Load Theory (germane vs extraneous load), Desirable Difficulties (productive struggle), Scaffolding (fading support), Over-Reliance (tool dependence), and Sociocultural Learning (general-purpose AI operates outside the ZPD by completing work for students). It is the bridge between assisted performance and genuine learning — the distinction between The Evidence Base on AI in K-12: A 2026 Review and the central question for AI Tutoring effectiveness.

Connected Concepts

Connected Articles

Embed this page

Copy the code below to embed a chromeless version of this page in a learning management system or other website. The embedded view hides the site header, navigation, and footer.