Research Article
ChatGPT in Education: An Effect in Search of a Cause
Synthesis: ChatGPT in Education: An Effect in Search of a Cause — A conceptual critique by Weidlich, Gašević, Drachsler, and Kirschner (2025) arguing that much early ChatGPT efficacy research repeats the classic media comparison fallacy: it compares an ill-defined "ChatGPT" treatment against opaque controls while measuring outcomes that are not durable learning. Using Deng et al.'s (2025) meta-analysis as a worked example, the authors revive methodological lessons from the Clark–Kozma media/methods debate to specify three "non-negotiable" conditions for interpretable causal effects — a precisely described treatment, a well-defined control group, and a valid measure of learning — and audit a subset of primary studies showing that only a small minority meet all three.
Key Findings
- Three non-negotiable conditions for interpretable effects. To claim any technology improves learning, researchers must (1) describe the exact nature of the experimental treatment so it can be replicated; (2) specify what the control group actually did — the operationalized counterfactual; and (3) use outcome measures that validly indicate durable learning rather than self-report or in-treatment performance.
- ChatGPT is a tool, not a method. Following Clark's (1983) "trucks don't change the groceries" analogy, learning does not follow from a general-purpose tool; effects arise only when the tool is embedded in a specific instructional method. Asking "does ChatGPT improve learning?" is a non sequitur — the meaningful question is whether a particular use of ChatGPT, in a particular instructional design, beats a particular comparison.
- Media-comparison confounding is pervasive. When new AI tools are introduced alongside new Pedagogies and Teaching Strategies, activity structures, or interface designs (Eronen's "fat-handed" interventions), the medium and method are confounded and effects become uninterpretable. This is a recurring pitfall documented across decades of research on instructional TV, computer-based instruction, distance education, AR/VR, and now AI.
- Deng et al.'s (2025) meta-analysis illustrates the problems. The audit of 19 academic-performance comparisons from Deng et al. found only 74% had a well-defined ChatGPT treatment, 42% a well-defined control group, and 53% an outcome that qualified as learning — leaving only 4 of 19 comparisons (21%) satisfying all three criteria.
- A stark effect-size anomaly. Deng et al. reported g = 0.7 for ChatGPT — larger than the 0.66 effect size of purpose-built Intelligent Tutoring Systems (Kulik & Fletcher 2016), even though ITS are engineered specifically for learning while ChatGPT was designed for entirely different purposes. Such a result signals that the "treatment" was a heterogeneous "secret sauce," not a coherent intervention.
- Valid learning measures are rare. Many outcomes were self-reported skills/motivation, or performance measured during treatment (confounding learning with task support), rather than unassisted, post-intervention measures of enduring change. A notable included study (Ahmed Moneus & Al-Wasy 2024) measured translation quality produced during human–ChatGPT collaboration — collaborative output, not learning — yet contributed a huge effect (g = 3.1).
- The lesson of "fast science." The rush to synthesize findings within two years of ChatGPT's launch (Deng et al. found 22 already-published ChatGPT-in-education reviews) produces research waste and premature causal claims. The authors call for a more deliberate research culture: let a richer literature accumulate, use stringent inclusion or detailed coding, and interpret meta-analytic effects cautiously.
Connected Concepts
- Generative AI
- Large Language Models (LLMs)
- Research Methods in AIED
- Meta-Analysis and Systematic Review
- Limitations in AIEd Research
- AI Ed Evaluation
- Learning Gains
- Intelligent Tutoring
- Assessment Validity
- RCT
Connected Articles
- Generative AI technologies and educational outcomes: a comprehensive meta-analysis comparing traditional and AI-driven approaches — A large meta-analysis of generative AI's effect on educational outcomes
- Critical AI Tutors: Empower or Enslave? — Critical limits of AI tutors and weak theory use
- Generative AI without guardrails can harm learning: Evidence from high school mathematics — Generative AI without guardrails can harm learning (PNAS 2025 RCT)
- Modeling AI Overreliance as a Complex Adaptive System — Overreliance on AI as a complex adaptive system
- How to Assess AI Literacy: Misalignment Between Self-Reported and Objective-Based Measures — Self-reported vs. performance-based AI literacy
Citation
Weidlich, J., Gašević, D., Drachsler, H., & Kirschner, P. A. (2025). ChatGPT in Education: An Effect in Search of a Cause. Journal of Computer Assisted Learning, 41(5), e70105.