Generative AI without guardrails can harm learning: Evidence from high school mathematics

Created: 2026-07-19 | Tags: generative-aiover-reliancestem-educationk-12rctlearning-gainsintelligent-tutoringscaffolding

Hamsa Bastani, Osbert Bastani, Alp Sungu, Haosen Ge, ร–zge Kurucu, Rehema Mushi (2025) โ€” Proceedings of the National Academy of Sciences (PNAS), 122(26). doi:10.1073/pnas.2422633122.

๐Ÿ“„ Full text (PNAS) โ€” RCT preregistered with the University of Pennsylvania IRB; anonymized data and code at github.com/obastani/GenAICanHarmLearning

Summary

This landmark field experiment is among the first randomized controlled trials to causally demonstrate that unguarded generative-AI tutoring can harm skill acquisition, not merely fail to help. Conducted with nearly 1,000 high-school math students across ~50 classes at a large school in Turkey (Fall 2023โ€“2024), the study compares three arms assigned at the classroom level: a control arm (textbook/notes only), GPT Base (a ChatGPT-like GPT-4 chat interface), and GPT Tutor (GPT-4 with teacher-designed guardrails โ€” hints instead of answers, plus the correct solution and common mistakes baked into the prompt). Each 90-minute session had three parts: a teacher lecture, an AI-assisted practice period (treatment only here), and an unassisted closed-book exam on conceptually similar problems.

The core finding: access to generative AI sharply improves performance during practice but, without guardrails, degrades learning once the tool is removed โ€” a tradeoff driven by students using the AI as a "crutch" to copy answers rather than to learn.

Key Findings

Performance vs. Learning Tradeoff (intention-to-treat)

Relative to control (practice mean 0.284; exam mean 0.321, normalized grades):

Arm Assisted practice (perf) Unassisted exam (learning)
GPT Base +48% (ฮฒ = 0.137\\) โˆ’17% (ฮฒ = โˆ’0.054\*)
GPT Tutor +127% (ฮฒ = 0.361\\) โ‰ˆ 0 (ฮฒ = โˆ’0.004, n.s.)

Mechanism: Students Use GPT Base as a Crutch

Two candidate explanations for the harm were tested โ€” (1) GPT Base's errors mislead students, and (2) students offload thinking by copying answers. The evidence strongly favors the second:

Students Don't Perceive the Harm

Students in the GPT Base arm performed worse on the exam but did not report learning or performing less; GPT Tutor users perceived they performed better than control even though exam scores were statistically indistinguishable. This perceived-vs-actual-learning mismatch parallels the "feeling of learning" literature and means self-report is an unreliable gauge of AI's learning impact.^[raw/papers/pnas-2025-guardrails-harm-learning.md]

Secondary Results

Design Lesson: What the Guardrails Did

GPT Tutor differed from GPT Base in two ways: (1) the prompt instructed it to give hints, not answers, and (2) it was seeded with teacher-authored problem-specific information (correct solution, common mistakes, feedback guidance) โ€” making its hints accurate and checkable. This labor-intensive prompt design is what neutralized the crutch effect. The authors note GPT Tutor remains passive (it doesn't proactively probe misconceptions) and call for combining pedagogical software tutors with generative AI, plus "co-pilot" models that assist human tutors rather than replace them.^[raw/papers/pnas-2025-guardrails-harm-learning.md]

Implications

Limitations (per authors)

Single topic (math), single high school in Turkey, Fall 2023 (early GPT-4 era), short-term outcomes only; writing and other subjects lack the objective grading used here. Generalizability to newer models and other contexts requires further study.^[raw/papers/pnas-2025-guardrails-harm-learning.md]

Related Pages

Citation

APA: Bastani, H., Bastani, O., Sungu, A., Ge, H., Kurucu, ร–., & Mushi, R. (2025). Generative AI without guardrails can harm learning: Evidence from high school mathematics. Proceedings of the National Academy of Sciences, 122(26). https://doi.org/10.1073/pnas.2422633122