On this page

Synthesis: This landmark field experiment is among the first randomized controlled trials to causally demonstrate that unguarded generative-AI tutoring can harm skill acquisition, not merely fail to help. Conducted with nearly 1,000 high-school math students across ~50 classes at a large school in Turkey (Fall 2023–2024), the study compares three arms assigned at the classroom level: a control arm (textbook/notes only), GPT Base (a ChatGPT-like GPT-4 chat interface), and GPT Tutor (GPT-4 with teacher-designed Guardrails — hints instead of answers, plus the correct solution and common mistakes baked into the prompt). Each 90-minute session had three parts: a teacher lecture, an AI-assisted practice period (treatment only here), and an unassisted closed-book exam on conceptually similar problems.

The core finding: access to generative AI sharply improves performance during practice but, without guardrails, degrades learning once the tool is removed — a tradeoff driven by students using the AI as a "crutch" to copy answers rather than to learn.

Performance vs. Learning Tradeoff (intention-to-treat)

Relative to control (practice mean 0.284; exam mean 0.321, normalized grades):

Arm Assisted practice (perf) Unassisted exam (learning)
GPT Base +48% (β = 0.137**) −17% (β = −0.054*)
GPT Tutor +127% (β = 0.361**) ≈ 0 (β = −0.004, n.s.)
  • GPT Base significantly improves practice performance but significantly worsens later unassisted exam performance — students who never had AI access actually outperformed them.
  • GPT Tutor's guardrails largely eliminate the learning harm (point estimate near zero), though they do not produce a positive learning effect either.

Mechanism: Students Use GPT Base as a Crutch

Two candidate explanations for the harm were tested — (1) GPT Base's errors mislead students, and (2) students offload thinking by copying answers. The evidence strongly favors the second:

  • Error analysis: When GPT Base's logical-error rate on a practice problem is higher, it hurts practice performance but shows no spillover to the corresponding unassisted exam problem — so students aren't being systematically misled into exam errors.
  • Engagement analysis: Students in the GPT Base arm send far fewer messages and overwhelmingly just "ask for the answer" or restate the question (superficial conversations dominate). In the GPT Tutor arm, a growing share of conversations are substantive (asking for help, attempting answers independently), and this improves within the very first session.
  • GPT Base answered correctly only 51% of the time on the 57 practice problems (42% logical errors, 8% arithmetic errors) — yet students still copy its outputs.

Students Don't Perceive the Harm

Students in the GPT Base arm performed worse on the exam but did not report learning or performing less; GPT Tutor users perceived they performed better than control even though exam scores were statistically indistinguishable. This perceived-vs-actual-learning mismatch parallels the "feeling of learning" literature and means self-report is an unreliable gauge of AI's learning impact.

Secondary Results

  • Skill-gap narrowing is temporary: Both AI arms reduced grade dispersion (HHI) during practice (biggest help to weakest students), but the effect does not persist on the unassisted exam.
  • Limited heterogeneity: Little evidence of differential effects by student ability, resources, or effort on exam performance.
  • Robustness: Intention-to-treat (including noncompliers), alternative specifications, and absenteeism checks all confirm the pattern.

Design Lesson: What the Guardrails Did

GPT Tutor differed from GPT Base in two ways: (1) the prompt instructed it to give hints, not answers, and (2) it was seeded with teacher-authored problem-specific information (correct solution, common mistakes, feedback guidance) — making its hints accurate and checkable. This labor-intensive prompt design is what neutralized the crutch effect. The authors note GPT Tutor remains passive (it doesn't proactively probe Misconceptions about AI) and call for combining pedagogical software tutors with generative AI, plus "co-pilot" models that assist human tutors rather than replace them.

What this means for practice

  • Learners. Demand hint-mode tutoring rather than answer-mode: GPT Base access left students 17% worse on the unassisted exam (beta = -0.054), while GPT Tutor's hint-not-answer guardrails reduced the harm to roughly zero (beta = -0.004) - the strongest causal, field-deployed evidence for the over-reliance mechanism.
  • Learners. Do not read your sense of learning as evidence of it: GPT Base users scored worse on the exam yet did not report learning less, and GPT Tutor users believed they outperformed the control arm when exam scores were statistically indistinguishable.
  • Learners. Ask for the correct solution and the common mistakes rather than the answer: GPT Tutor's hints were accurate because they were seeded with teacher-authored, problem-specific information, while GPT Base answered only 51% of the 57 practice problems correctly.
  • Learners. Practice with the tool removed before trusting what it taught: assisted practice performance rose 48% (GPT Base) and 127% (GPT Tutor), yet exam gains did not follow and the narrowing of the grade gap during practice did not persist.

Limitations

  • One topic, one site, one snapshot: the field experiment ran on mathematics in a single high school in Turkey across Fall 2023-2024, in the early GPT-4 era, so replication with newer models and other contexts is required.
  • Outcomes are short-term only: the unassisted closed-book exam came after each 90-minute session, and the authors note that writing and other subjects lack the objective grading used here.
  • Study arms were assigned at the classroom level rather than individually, so treatment is clustered by class even though intention-to-treat, alternative specifications and absenteeism checks support the pattern.
  • Perceptions of learning were miscalibrated against measured exam performance in both AI arms, so self-reported learning or performance cannot serve as a gauge of the intervention's effect.

Citation

Bastani, H., Bastani, O., Sungu, A., Ge, H., Kurucu, Ö., & Mushi, R. (2025). Generative AI without guardrails can harm learning: Evidence from high school mathematics. Proceedings of the National Academy of Sciences, 122(26).

Embed this page

Copy the code below to embed a chromeless version of this page in a learning management system or other website. The embedded view hides the site header, navigation, and footer.