The null average performance effect masks strong offsetting heterogeneity β and the exam had severe ceiling compression (control mean 89.2/100, 47% β₯ 95), which also limits power. The belief reversal is striking: it contradicts "familiarity breeds acceptance" and suggests an arc from initial awe at AI's instant responses to awareness of its unintended effects.
Sungu, Lira & Duckworth (2026) ran one of the first large-scale RCTs of a teacher-facing generative AI tool and found it can harm students: providing teachers an AI teaching assistant reduced student intrinsic motivation by 0.11 SD and β among lower-performing teachers β cut student achievement by 0.13 SD. The pattern is a principalβagent problem: teachers (agents) gain labor savings from AI delegation while students (principals) bear the cost of displaced relational teaching and scaffolding.
The experiment
538 teachers across 24 Turkish K-12 schools randomized at school-department level; analytical sample 193 teachers / 2,816 students / 14,198 student-course observationsTreatment: custom GPT-4o chatbot with Turkish Ministry of Education curriculum database + 1-hour training (one arm added weekly usage-stat reminders); control = business-as-usualPre-registered; ITT; semester-length (spring 2025)Results
| Outcome | Average effect | Heterogeneity |
|---|
| Student intrinsic motivation | β0.111 SD (p=.015) | Heavy baseline AI users: β0.182 (p=.015); light users: β0.052 (ns) |
| Student confidence | β0.090 SD (p=.097) | Lower-performing teachers: β0.183 (p=.012); higher: β0.022 (ns) |
| Academic performance | β0.019 SD (ns, ceiling-compressed) | Below-median teachers' students: β0.129 (p=.005); above-median: +0.054 (ns) |
| Teacher beliefs about AI's effect on learning | +0.126 SD (ns) | Heavy prior users became more pessimistic (β0.379); light users more optimistic (+0.458) |
The null average performance effect masks strong offsetting heterogeneity β and the exam had severe ceiling compression (control mean 89.2/100, 47% β₯ 95), which also limits power. The belief reversal is striking: it contradicts "familiarity breeds acceptance" and suggests an arc from initial awe at AI's instant responses to awareness of its unintended effects.
Why the harm happens: usage patterns
66% of teacher conversations were teaching-material production (lecture prep 32%, homework/exam 22%, syllabus 9%); only 16% instructional support; 18% generalShallow use: median 2 prompts, mean 4.7 messages per session β teachers accepted outputs with minimal iterationInterpretation: task delegation, not pedagogical collaboration β the tool was a generator of finished artifacts rather than an iterative partner, limiting the pedagogical reflection that separates augmentation from substitutionConnected Concepts
Generative AIK 12Student ExperienceTeacher AI CompetencyTeacher RoleRAGConnected Articles
Beyond Detection Authentic Assessment AI 2025 β Beyond Detection: redesigning authentic assessment in an AI-mediated worldCare Full Feedback GenAI β The care-full craft of feedback in an age of generative AIGenAI Expertise Pathways Sysadmin β Unanticipated Effects of Generative AI on Expertise Pathways and Performance Perception in System AdministrationOecd Digital Education Outlook 2026 β OECD Digital Education Outlook 2026Aaai2026 Prompting Literacy K12 β Learning to Use AI for Learning: Teaching Responsible Use of AI Chatbot to K-12 Students Through an AI Literacy ModuleAcademiclaw Student Agent Benchmark β AcademiClaw: When Students Set Challenges for AI AgentsAccess Not Enough AI Tutoring 2026 β Access is Not Enough: Human Support Improves Engagement with AI TutoringAdapt Adaptive Lesson Plan Transformer β AdaPT: Adaptive Lesson Plan Transformer for Cross-Regional and Differentiated InstructionAdaptive Pretesting Retention β Do Gains from Generative AI-Enabled Adaptive Pretesting Persist? Evidence from a Retention StudyAffective Text Wearable Student Health β A Formative Study of Brief Affective Text as a Complement to Wearable Sensing for Longitudinal Student Health MonitoringAgency Gap AI Writing β The agency gap in AI-supported writing: how reactive and proactive agent designs shape multimodal reasoningAgent Voice Accents K12 Group Learning β Exploring How Agent Voice Accents Shape Human-AI Collaboration in K-12 Group LearningAgentic AI Education Scoping Review β Agentic AI in Education: A Scoping Review of Research Landscape, Capabilities, and the Frontier Agent ParadigmAgentic Education Coding β Agentic Education with AI Coding AssistantsAgentic Literacy Debt β Agentic Literacy Debt: A Structural Problem the AI Literacy Field Has Not Yet NamedAgents That Teach Incidental Learning β Agents That Teach: Designing Incidental Learning Back into AI-Assisted Software DevelopmentAgreement Not Quality LLM Coding Verification β Agreement Is Not Quality: Blind Expert Verification of Human and LLM Qualitative Coding When Human Consensus Is Not G...AI Adult Learning Design β Guidelines for Designing AI Technologies to Support Adult LearningAI Assessment Human Tutors β AI-Driven Assessment of Human Tutors: Linking Training Performance to Real-Life PracticeAI Assessment Scale Reform β A bit of chaos and madness": The AI Assessment Scale and the work of assessment reformAI Assistance Discretionary Feedback β AI Assistance for Discretionary Work: Increasing Feedback Provision in Higher EducationAI Assisted Learning Modes Eeg β An exploratory behavioral and electroencephalographic study of artificial intelligence-assisted learning modes in hig...AI Assisted Se Curriculum Syllabus Analysis 2026 β Mapping the Emerging Curriculum for AI-Assisted Software Engineering via Syllabus AnalysisAI Assisted Writing Research Teams β Smaller, Younger, and More Impactful: How AI-Assisted Writing Transforms Research TeamsAI Availability Student Motivation β Why Put in This Much Effort?": How AI Availability Shapes Studentsβ Motivation in Introductory ProgrammingCitation
Sungu, Lira & Duckworth (2026). Generative AI Can Harm Teaching