On this page

Synthesis: Wagner-Kobayashi (2026) asks whether ChatGPT used inside a writing-to-learn task helps students learn the material they write about. Because composing with elaboration prompts is theorized to build knowledge and ChatGPT supplies examples and connections on demand, H1 predicted a steeper gain for the assisted group. In a laboratory experiment, 35 German-speaking psychology students (n = 16 with ChatGPT, n = 19 without, after five exclusions from forty) spent 20 minutes elaborating an input text, then took an identical knowledge test. The group × time interaction was significant (F(70) = 5.889, p = .018) but pointed the other way: the no-ChatGPT group gained more (posttest M = 8.29 vs 6.94), with b = −1.450 against the hypothesis. Motivation (p = .009) and interest (p = .017) interacted with group, so students low in motivation or topic interest were disadvantaged. Takeaway: untrained, time-pressured ChatGPT use did not benefit learning and may have harmed it.

Key Findings

  1. The hypothesis was falsified in direction, not for absence of an effect. The predicted group × time interaction was significant (F(70) = 5.889, p = .018), but b = −1.450 showed that participants writing without ChatGPT had the steeper slope.
  2. The means make the reversal concrete. The experimental group moved from M = 1.5 to M = 6.94 on the knowledge-test sum, the control group from M = 0.79 to M = 8.29 — a larger gain from a lower start.
  3. Motivation and interest mattered only for ChatGPT users. Both entered significant group × covariable interactions (F(70) = 7.183, p = .009 and F(70) = 6.027, p = .017), while adding motivation made group × time non-significant (F(70) = 1.187, p = .280).
  4. Nothing else predicted the effect, and few elaborations were generated. Self-Efficacy for self-regulation of academic writing, concentration, examples, connections and writing belief were non-significant, while the ChatGPT group produced only 2.25 of its 2.94 examples and 0.88 of 1.81 connections with the tool — a utilization deficiency (Miller, 2000).
  5. The finding is statistically fragile and the authors say so. Residuals deviated from normality (Shapiro-Wilk W = 0.96, p = .004), variance assumptions failed in two models, and the level-two sample fell short of N = 50.

Why a benefit was expected: writing-to-learn theory

Writing is a mode of learning: Emig's (1977) "writing as a learning mode" and the meta-analytic evidence of Graham et al. (2020) and van Dijk et al. (2022) ground the prediction that composing about content improves knowledge of it. The mechanism targeted is deliberate elaboration — generating examples and connections — and cognitive accounts of writing hold that this work converts reading into durable knowledge. ChatGPT appeared to fit that mechanism almost perfectly by generating examples and comparisons time-efficiently, so large language models should amplify a writing-to-learn task. Against this they set the "double-edged sword" framing — integrity, accuracy, Privacy, erosion of higher-order thinking, worse inequality for lower-skilled Learners — and Kosmyna et al.'s (2025) claim that support tools "restructure not only task performance but also the underlying cognitive architecture".

Design, task, and how ChatGPT use was logged

A between-subjects laboratory experiment ran in Frankfurt from late January to mid-March 2024, advertised as "digital writing-to-learn" without mentioning ChatGPT so the control group would not adjust effort. Of the 40 recruited (20 per condition), five were excluded — three who never used ChatGPT for connections or examples, one who wrote the full text with it, and one outlier who scored zero on both tests and left a 20-minute task after 12 minutes — leaving 35 (n = 16 experimental, n = 19 control). After a demographic questionnaire and pretest, participants wrote for 20 minutes from a modified Knaus (2011) extract with instructions to elaborate by generating examples and connections; the experimental group had ChatGPT (gpt-3.5-turbo-1106, via the university's IP registration) on a second tab, and submitting its complete output as their own text was forbidden. The same ten-question open-answer test — summed scores operationalizing learning performance — was given again without ChatGPT, and analysis used repeated-measures hierarchical linear models.

Results: the benefit that did not appear

Both groups learned — time was significant (p < .001) in every model — but the ChatGPT group learned less. The confirmatory model found the hypothesized interaction, F(70) = 5.889, p = .018, and then contradicted it: with control at posttest as reference, b = −1.450, so the steeper slope belonged to the unaided writers. Adding motivation made group × time non-significant (F(70) = 1.187, p = .280) while motivation (F(70) = 7.451, p = .008) and its interaction with group (F(70) = 7.183, p = .009; b = .618, p = .031) became significant. Interest repeated that pattern more weakly (group × interest F(70) = 6.027, p = .017), and the two correlated at r = .805. Motivation and interest protected only the group whose task carried the tool's extra demands: drawing on volition research (Heckhausen, 1991), the authors suggest motivation shielded ChatGPT users from added load.

Participants reported that time pressure shifted their focus to quantity over depth of engagement; the authors suggest ChatGPT users asked the same questions repeatedly and pasted answers instead of seeking their own examples and connections to prior knowledge, reading the low elaboration counts as a utilization deficiency, the lag between having a new instrument and knowing how to deploy it. They also raise cognitive overload: integrating generated content into a coherent essay may have displaced processing of the content itself. The findings converge with Gayed et al.'s (2022) null result for AI writing assistance.

Effort, ownership, and what this means for cognitive offloading

The study is a causal test of an offloading prediction inside a single writing task, and its negative result is narrow. What leaked away in the ChatGPT condition was elaborative effort — generating one's own examples and connections — the activity writing-to-learn credits with the learning, so the displacement lands on the target skill. The utilization-deficiency reading cuts the other way: the authors do not claim ChatGPT cannot support critical thinking, only that it did not here, untrained and under a 20-minute limit, so the tool may be mis-deployed or inert. They recommend Scaffolding — training and clear deployment instruction — rather than prohibition, plus caution where intrinsic motivation and topic interest vary widely. Two features limit this reading: the ChatGPT group was the more motivated one yet gained less, and outsourcing the whole text was forbidden, so the harm occurred where the tool was meant to assist elaboration — the setup most Pedagogies and Teaching Strategies recommends.

What this means for practice

  • Instructors. Teach the tool before you assign it: the experimental group produced only 2.25 of 2.94 examples and 0.88 of 1.81 connections with ChatGPT, so training could have mitigated it.
  • Instructors. Protect elaboration time and require a visible product of student-generated elaboration, the activity writing-to-learn credits with the learning and the assisted condition displaced; participants focused on the word requirement, and ChatGPT users pasted answers rather than seeking their own examples.
  • Instructors. Front-load motivation and topic-interest support whenever a task carries an AI tool: higher motivation or interest predicted a steeper gain only inside the ChatGPT group (p = .009 and p = .017), which started higher on motivation (M = 8.56 vs 7.26) yet gained less.

Limitations

  • The sample of 35 is small and psychology-only, uneven across conditions, and the authors call the randomization suboptimal for lacking counterbalancing or participant matching.
  • Several RM-HLM assumptions were violated — non-normal residuals, heterogeneity of variance in two models and a level-two N below 50 — which underestimates standard errors, though floor effects at pretest plausibly explain it.
  • Ecological validity is questionable (an artificial exam-like setting, unfamiliar desktops, one 20-minute episode), the manipulation was imperfect — 3 of 20 ChatGPT-group participants never used the tool — and the unvalidated test mainly captures lower levels of Bloom's taxonomy.

Connected Concepts

  • Cognitive Offloading — the prediction the study tests: delegation of elaborative work to a model
  • Critical Thinking — self-generated examples and connections as the thinking the task required
  • Writing — writing-to-learn as the instructional frame the intervention sits inside
  • Metacognition — deliberate elaboration as the mechanism, and the effort participants did not spend
  • Generative AI — ChatGPT 3.5 as the tool in the experimental condition
  • Large Language Models (LLMs) — the model class whose time efficiency was expected to amplify a writing-to-learn effect
  • Conversational AI — the chat interface through which participants generated examples
  • Motivation — the pre-existing group difference that absorbed variance attributed to ChatGPT use
  • Self-Efficacy — self-efficacy for self-regulation of academic writing, measured and non-significant
  • Transfer of Learning — what a single 20-minute writing task can and cannot show about durable learning
  • Desirable Difficulties — the effortful conditions the ChatGPT condition removed
  • Retrieval, Spacing and Interleaving — self-generated elaboration contrasted with fluent, model-supplied content
  • Cognitive Psychology — the learning and memory findings the hypothesis was built on

Connected Articles

Citation

Wagner-Kobayashi, E. L. (2026). ChatGPT making our minds dull? The cognitive impact of using ChatGPT in the writing process. PsyArXiv Preprints.

Embed this page

Copy the code below to embed a chromeless version of this page in a learning management system or other website. The embedded view hides the site header, navigation, and footer.