On this page

OECD flagship report synthesizing empirical evidence and expert insights on generative AI and AI in education. Central finding: general-purpose AI chatbots improve task performance but produce no durable learning gains; purpose-built educational GenAI, co-designed with teachers, is the path to sustained improvement.

The Core Finding: Performance Is Not Learning

The report's most consequential finding is that general-purpose GenAI tools (ChatGPT, Gemini, Claude) improve the quality of student outputs on assignments but this advantage disappears and sometimes reverses in exams when AI access is removed. This misalignment between task performance and genuine learning is the report's central argument for why purpose-built educational AI is necessary. When students offload cognitive tasks to chatbots, metacognitive engagement drops — the mental processes that turn answers into understanding are short-circuited. See Distinguishing performance gains from learning when using generative AI and Over-Reliance.

Educational GenAI: What Works

Hybrid systems that combine GenAI with explicit pedagogical models show more promise than general-purpose chatbots. Examples cited include:

  • Socratic Playground (SPL): Uses Socratic questioning to develop subject knowledge, critical thinking and reflection rather than providing direct answers
  • Khanmigo: Withholds answers and guides reasoning through questioning
  • JeepyTA: AI teaching assistant in university contexts rated comparable to human TAs in clarity and accuracy
  • Tutor Copilot: Mobilizes less-qualified tutors effectively through AI support

The report draws a sharp line: GenAI tools "designed or used with an intentional pedagogical purpose" produce sustained learning improvements; tools used as answer-dispensing shortcuts do not. See AI Tutoring.

Tutoring: The 9-Percentage-Point Effect

A 9-percentage-point increase in student pass rates when low-experience tutors used AI support, with smaller gains for more experienced tutors. Secondary science teachers in England saw a 31% reduction in time spent on lesson and resource planning. This supports an augmentation model where AI boosts the least experienced practitioners the most. See Improving Hybrid Human-AI Tutoring by Differentiating Human Tutor Roles Based on Student Needs.

Teacher Agency: Three Paradigms

The report's conceptual framework (Ch.7) proposes three paradigms for teacher-AI interaction:

  1. Replacement — AI takes over tasks; risk of teacher deskilling
  2. Complementarity — Human judgment paired with machine efficiency
  3. Augmentation — Teachers and AI work in tandem, critiquing and refining each other's outputs (recommended)

The augmentation paradigm preserves professional judgment while achieving the greatest instructional quality gains. See Generative AI Can Harm Teaching.

Purpose-Built vs General-Purpose

Chapter 8 makes the case for purpose-built educational GenAI systems co-created with teachers and students. These tools would give teachers control over AI behavior — including setting the level of "hallucinations" — and enable monitoring of student-AI interactions. The tools should align with specific curricula rather than being generic, and maintain teacher autonomy over course design and enactment.

Collaborative Learning and Creativity

GenAI supports Collaborative Learning in four roles: information hub, personalized material generator, teacher feedback provider, and peer contributor. Studies find small-to-medium improvements in subject learning and large ones in critical thinking and teamwork (Ch.4). For Creativity, GenAI works best when used "slowly" for iterative exploration and reflection, not for instant content generation (Ch.5).

System-Level and Assessment Applications

At the institutional level, GenAI enables: curriculum mapping between courses/programs, admissions and career guidance analytics, standardized assessment item generation, interactive writing and speaking assessments, and synthetic datasets for education research (Chs. 11–13).

Policy Recommendations

Four pillars: (1) human-centered teaching and learning with GenAI; (2) investment in educational GenAI R&D grounded in learning science; (3) enabling policy environment for trustworthy GenAI (privacy, safety, bias testing, transparency); (4) equitable digital infrastructure including offline small language models for low-connectivity settings.

Equity: AI Unplugged

A large-scale experiment in rural Brazil (Ch.6) demonstrated that even with intermittent connectivity and minimal equipment, AI could provide feedback and guidance. Small language models running offline on mobile devices are identified as a promising avenue for bridging digital divides.

What this means for practice

  • Instructors. Build GenAI into tasks with an explicit pedagogical purpose rather than leaving it as an answer service: general-purpose chatbots improved the quality of student output, but the advantage disappeared and sometimes reversed once AI access was removed.
  • Faculty developers. Train staff toward augmentation — teachers and AI critiquing and refining each other's output — rather than replacement, which the report associates with deskilling.
  • Administrators. Target AI support where it moves outcomes most: low-experience tutors gained 9 percentage points in student pass rates with AI support, with smaller gains for experienced tutors, and secondary science teachers in England cut lesson and resource planning time by 31 percent.
  • Administrators. Assume students will use GenAI regardless and redesign accordingly — assignments that cannot be completed directly by a chatbot, oral defense in lab time, and conceptual paper exams produced comparable outcomes for groups with and without GenAI access in one reported course redesign.
  • Researchers. Measure retention, not task performance: in a randomized trial with about 1,600 students, AI-supported gains in peer feedback quality were not sustained after the tool was withdrawn.

Limitations

  • The Outlook is a secondary synthesis of existing empirical studies and expert input rather than new primary data collection, so its conclusions inherit the designs, samples, and settings of the studies reviewed.
  • Several headline figures come from single studies in specific settings — the 9-percentage-point pass-rate gain and the 31 percent planning-time reduction — and are reported without replication across systems.
  • The TALIS-based finding that 37 percent of teachers use GenAI for work-related tasks is flagged in the report itself as carrying a higher risk of non-response bias and should be interpreted with caution.
  • Evidence on synergy is mixed: a meta-analysis of 106 experimental studies of human-AI collaboration found that, on average, human-AI combinations performed worse than the best of either humans or AI alone, particularly on decision-making tasks.

Citation

OECD (2026). OECD Digital Education Outlook 2026: Exploring Effective Uses of Generative AI in Education. OECD Publishing, Paris.

Embed this page

Copy the code below to embed a chromeless version of this page in a learning management system or other website. The embedded view hides the site header, navigation, and footer.