π Full text: SSRN 7007339 Β· local
Sungu, Lira & Duckworth (2026) ran one of the first large-scale RCTs of a teacher-facing generative AI tool and found it can harm students: providing teachers an AI teaching assistant reduced student intrinsic motivation by 0.11 SD and β among lower-performing teachers β cut student achievement by 0.13 SD. The pattern is a principalβagent problem: teachers (agents) gain labor savings from AI delegation while students (principals) bear the cost of displaced relational teaching and scaffolding.
The experiment
- 538 teachers across 24 Turkish K-12 schools randomized at school-department level; analytical sample 193 teachers / 2,816 students / 14,198 student-course observations
- Treatment: custom GPT-4o chatbot with Turkish Ministry of Education curriculum database + 1-hour training (one arm added weekly usage-stat reminders); control = business-as-usual
- Pre-registered; ITT; semester-length (spring 2025)
Results
| Outcome | Average effect | Heterogeneity |
|---|---|---|
| Student intrinsic motivation | β0.111 SD (p=.015) | Heavy baseline AI users: β0.182 (p=.015); light users: β0.052 (ns) |
| Student confidence | β0.090 SD (p=.097) | Lower-performing teachers: β0.183 (p=.012); higher: β0.022 (ns) |
| Academic performance | β0.019 SD (ns, ceiling-compressed) | Below-median teachers' students: β0.129 (p=.005); above-median: +0.054 (ns) |
| Teacher beliefs about AI's effect on learning | +0.126 SD (ns) | Heavy prior users became more pessimistic (β0.379); light users more optimistic (+0.458) |
The null average performance effect masks strong offsetting heterogeneity β and the exam had severe ceiling compression (control mean 89.2/100, 47% β₯ 95), which also limits power. The belief reversal is striking: it contradicts "familiarity breeds acceptance" and suggests an arc from initial awe at AI's instant responses to awareness of its unintended effects.
Why the harm happens: usage patterns
- 66% of teacher conversations were teaching-material production (lecture prep 32%, homework/exam 22%, syllabus 9%); only 16% instructional support; 18% general
- Shallow use: median 2 prompts, mean 4.7 messages per session β teachers accepted outputs with minimal iteration
- Interpretation: task delegation, not pedagogical collaboration β the tool was a generator of finished artifacts rather than an iterative partner, limiting the pedagogical reflection that separates augmentation from substitution
Connections to the wiki
- Direct causal evidence for the cognitive-offloading concern applied to teachers rather than students, and for over-reliance at the instructor level
- A counterweight to faculty-development narratives that AI support tools straightforwardly improve teaching β effectiveness depends on how the tool is used (teacher-ai-competency)
- The motivational harm connects to student-experience and the relational critiques in care-full-feedback-genai (feedback as "matters of care" is displaced when AI mediates material production)
- Skill-substitution channel mirrors the genai-expertise-pathways-sysadmin finding that GenAI compresses expertise pathways and resets performance expectations
- Complements the design-not-detection agenda of beyond-detection-authentic-assessment-ai-2025: teacher-facing AI needs the same design scrutiny as assessment-facing AI
Related Pages
- faculty-development β teacher-facing AI adoption and its unintended effects
- teacher-ai-competency β what teachers need to use AI as augmentation, not substitution
- cognitive-offloading β the mechanism behind degraded pedagogical reasoning
- over-reliance β shallow, accept-output AI use
- student-experience β motivational and confidence harm to students
- generative-ai β teacher-side deployment effects
- k-12 β K-12 context of the field experiment
- genai-expertise-pathways-sysadmin β parallel expertise-compression finding
- care-full-feedback-genai β relational teaching displaced by AI mediation
- teacher-role β principalβagent tension in AI-assisted teaching
Sources
- Sungu, A., Lira, B., & Duckworth, A. L. (2026). Generative AI Can Harm Teaching. SSRN Working Paper 7007339. SSRN