On this page

Synthesis: Ratajczyk, Dymarska, Matłoka, Tomczyk and Wiącek ask the question most genAI-in-education research dodges: not whether ChatGPT helps while the window is open, but what is left of a learner's reasoning once it closes. Their preregistered laboratory experiment (final N = 193) had participants work through a fictional 15-entity relational network across three stages — a chunk-based reasoning task, a harder network-level reasoning task built on the same material, and a memory test — with ChatGPT available only in Stage 1, crossed against a performance-incentive manipulation. Access raised Stage 1 accuracy (77.9% vs 72.7%, ηp² = .030) and lowered perceived difficulty (4.17 vs 4.71 on a seven-point scale). In the unaided Stage 2 there was no aggregate accuracy difference, but item-level modelling found prior genAI use associated with roughly 32% lower odds of a correct answer (OR = 0.68). The starker result was motivational: incentives lengthened time on difficult Stage 2 items for participants who had worked alone but not for those who had used the tool. The authors read this as cognitive offloading weakening effort regulation rather than reasoning capacity itself.

Key Findings

  1. A five-condition preregistered lab experiment, not a classroom study. The design was 2 (genAI access: available vs unavailable) × 2 (performance incentive: present vs absent) between subjects, plus a control condition that worked on a structurally similar but unrelated network in Stage 1. Final sample 193 (39, 39, 40, 39, and 36 per condition), from 198 tested; one participant was excluded for non-compliance and four for failing attention checks. Mean age 25.1 years (SD = 7.0); 105 had completed secondary and 88 higher education.
  2. GenAI helped immediately, and the incentives did not. Stage 1 accuracy was 75.3% (SD = 14.9%) overall. Participants with ChatGPT scored 77.85% (SD = 14.36) against 72.72% (SD = 15.04) for those working independently, F(1, 153) = 4.71, p = .030, ηp² = .030. Incentives produced no accuracy gain (F(1, 153) = 1.64, p = .207) but did extend Stage 1 completion time, 1262 s vs 1077 s, F(1, 153) = 9.82, p = .002, ηp² = .060.
  3. The tool made the same task feel easier. Perceived Stage 1 difficulty was 4.17 (SD = 1.48) in the genAI condition against 4.71 (SD = 1.42) without it, F(1, 153) = 5.49, p = .020, ηp² = .035, on a seven-point scale. This perceived-difficulty reduction is the immediate cognitive-load benefit the paper is careful not to dismiss.
  4. Unassisted reasoning slipped, but only at the item level. Stage 2 overall accuracy was 61.6% (SD = 17.2%), and the preregistered aggregate test found no main effect of prior genAI access (F(1, 153) = 1.84, p = .177, ηp² = .011). In an exploratory mixed-effects logistic regression, prior genAI access within the no-incentive reference condition was associated with lower odds of a correct Stage 2 answer, b = −0.392, SE = 0.195, z = −2.01, p = .045, OR = 0.68 — about 32% lower odds. The genAI × incentive interaction was not significant (p = .141), so this is a conditional effect.
  5. Effort regulation was where the difference concentrated. On Stage 2 completion time the incentive main effect was significant (874.6 s vs 751.5 s, F(1, 153) = 8.44, p = .004, ηp² = .052) and so was the genAI × incentive interaction, F(1, 153) = 4.38, p = .038, ηp² = .028. Without prior genAI access, incentives stretched completion time from 707.4 s to 941.2 s (t(76) = 3.72, p_adj = .002); within the genAI condition there was no such adjustment (t(72) = 0.56, p_adj = .58).
  6. The effect was specific to hard problems. A three-way interaction of prior genAI access, incentive and question difficulty was significant, b = −0.51, SE = 0.24, t(1869) = −2.08, p = .038, and held only at high difficulty (b = −0.497, p = .004). Among incentivised participants on difficult items, the no-genAI group took 126.4 s against 84.7 s for the genAI group (p = .008). Stage 2 perceived difficulty itself did not differ across conditions (F(1, 153) = 0.21, p = .646), so this was not a case of the tool users finding the unassisted task unbearable.
  7. Memory showed nothing. Stage 3 performance was 57.5% (SD = 17.6%) with no main effect of prior genAI access (F(1, 153) = 0.095, p = .758) and no incentive effect (ηp² < .001). The one Stage 3 difference was meta-level: among participants with prior genAI access, those incentivised rated the memory test harder, 6.10 vs 5.79, t(72.07) = −2.84, p_adj = .035, d = −0.64.
  8. Who offloaded was not predicted by working memory or incentives. Incentives did not change the number of ChatGPT queries (U = 664.0, p = .334, r_rb = −.126), and working memory capacity did not predict query count (F(1, 76) = 0.01, p = .903, R² < .001). Almost 40% of participants with access either never used ChatGPT or used it only occasionally, and those participants held markedly less favourable attitudes to AI (4.42 vs 6.10 and 6.52 on AIAS-4, F(2, 72) = 6.27, p = .003, ηp² = .148).
  9. Offloading came in different shapes. Of 78 codable conversations (Cohen's κ = 0.85), 31 showed no or occasional use, 28 mostly copied questions into the tool, 16 used it exploratorily by proposing answers, asking for justification or challenging the model, and 3 used it for other purposes. Stage 1 accuracy rose across the groups (74.4%, 81.7%, 82.3%, F(2, 72) = 3.15, p = .049, ηp² = .080) without any pairwise difference reaching significance, while exploratory users spent far longer on the task (1473 s vs ~1140 s, p_adj = .041 for both comparisons). Queries correlated with Stage 1 accuracy (rs = .376, p < .001), and harder items drew more participants to the tool (r(16) = .607, p = .008).

The three-stage task and what each stage isolates

The stimulus was a network of 15 fictional entities — creatures in the main version, celestial bodies in the control — described by 82 unique logical propositions, deliberately too large to hold in working memory at once. Stage 1 (chunk-based reasoning) presented subsets of four to eight statements followed by inference questions with three response options: yes, no, or impossible to determine. Chunking follows cognitive load theory's element interactivity logic: the point of Stage 1 is to create an opportunity to build a structured representation of the network, not merely to answer ten questions. Stage 2 (network-level reasoning) handed participants 20 paper information points containing 48 propositions that had appeared in Stage 1 and asked 13 questions requiring integration across them. Every question was answerable from those points, so Stage 2 measures the ability to reconstruct and apply relational structure unaided. Stage 3 was a 10-item, four-option memory test with no access to any of the material.

The three-stage design attributes each stage to a different theoretical object, and that is what makes it informative. Stage 1 is where genAI access and incentives are manipulated, so it captures assisted performance and cognitive load. Stage 2 removes the tool and asks whether the relational knowledge needed for later inference was actually constructed. Stage 3 asks whether the individual propositions were retained. The control condition, which completed a comparable Stage 1 on a different network, is the clever part: it separates "offloading hurt knowledge construction" from "the tool just made people tired or less practised", because control participants carried the same task demands without acquiring any transferable relational knowledge for Stage 2. Notably, the control group's Stage 2 accuracy (57.3%) was numerically below the no-genAI baseline (64.1%) but not significantly so (t(73) = 1.79, p = .078, d = 0.41), while taking significantly longer to finish (U = 435.00, p = .005, r_rb = −.380). Participants who lacked prior network knowledge compensated with time — a benchmark against which the genAI group's failure to add time looks pointed.

How effort regulation was operationalised

Effort appears in the study in three registers. Behaviourally, it is completion time on Stage 1 and Stage 2, analysed at the item level so that the interaction with question difficulty can be estimated; the preregistered hypotheses treat faster completion as ambiguous, since it could signal either fluency or withdrawal, which is why the authors rely on how time responds to incentives rather than on absolute speed. Subjectively, it is the Effort/Importance subscale of the Intrinsic Motivation Inventory, translated into Polish with back-translation verified by two bilinguals. Meta-cognitively, it is a single seven-point perceived-difficulty rating for each stage, and the AIAS-4 attitude scale plus self-reported genAI usage frequency.

The independent variable that makes effort visible is the incentive manipulation: participants were told the top-performing third would receive a bonus, that Polish results would be compared with other European countries, and that individual scores could be released. The logic comes from motivation intensity theory, where effort rises with task difficulty as long as success still looks possible and the required effort stays within what the importance of success justifies. Crucially, the effect of incentives was to buy time, not accuracy — including in Stage 1, where incentives lengthened completion but produced no accuracy difference. That keeps completion time interpretable as effort allocation, and it means the interesting outcome is the slope of time against difficulty and incentive, not a performance score. Self-reported effort showed the same directional pattern as the behavioural data: the genAI × incentive interaction was significant (F(1, 153) = 4.34, p = .039, ηp² = .028), with numerically higher declared effort under incentive for those who had worked without genAI (U = 1003.5, p = .088, Hedges' g = 0.40) and a reversed pattern among prior tool users, though no simple comparison survived correction.

Assisted performance, then unassisted reasoning: what the results support

Read against the paper's preregistered hypotheses, the headline is negative. H1a and H1b — that prior genAI access would improve, or impair, Stage 2 accuracy — were both unsupported at the aggregate level. H3 (better memory without assistance) was unsupported. H6 and H7, on the determinants of offloading, were unsupported. What survived was an exploratory item-level effect in the direction of impairment, plus the pre-registered incentive effect on time (H5) and its dependence on prior genAI access.

The item-level picture is coherent even where the aggregate picture is flat. Mixed-effects models that treat participants and items as random effects find that prior tool users had roughly a third lower odds of answering an individual Stage 2 question correctly within the no-incentive condition, and that their time investment on difficult questions did not rise when incentives made success more valuable. The three-way interaction is the cleanest expression: without genAI, incentives added time to hard items (81.8 s to 126.4 s, p = .004); with prior genAI, time barely moved (p = .208). Two dissociations support reading this as regulation rather than ability. First, Stage 2 difficulty was perceived identically across conditions, so the tool users were not overwhelmed by the unassisted task. Second, memory was untouched, suggesting that individual facts survived even where their relational use did not — consistent with the authors' distinction between remembering propositions and being able to derive new conclusions from them. The Stage 2 control comparison adds a third: the participants who genuinely lacked prior knowledge responded by spending more time, and it did not rescue their accuracy either. Time was the currency everyone had, and prior genAI users were the only group that did not spend it when spending it would have helped.

The mechanism: effort withdrawal, calibration, or reliance habit

The paper's preferred reading is a change in effort regulation, and it parks two mechanisms side by side. The first is motivation intensity theory's potential motivation: prior genAI success may have lowered how much personal effort participants judged this task to warrant, so hard questions hit that subjective ceiling sooner. The second is expected value of control: Stage 1 taught participants an inaccurate contingency between personal effort and success, and when the tool vanished they failed to recalibrate, allocating too little control to difficult items. The self-report pattern leans slightly toward the first — people appeared at least partly aware of investing less — but since no simple effect survived correction, the authors decline to choose and note the two may operate together. They also rule out a shallow alternative: participants did not simply forget the incentive instructions, since incentives lengthened Stage 1 time even with the tool present.

This sits in a specific relation to neighbouring accounts. It is not the "cognitive fatigue" story, because the control condition carried similar load without the same result. It is not straightforward knowledge loss, because memory was unaffected. It is closest to Fan et al.'s metacognitive laziness, Liu et al.'s metacognitive decay (where assisted performance resets expectations of how effortful a task should be) and Wu et al.'s finding that motivation rises with genAI and then drops more steeply when it is withdrawn; Kosmyna et al.'s weaker neural engagement and poorer recall in genAI essay writers is the neuroimaging analogue. The distinctive contribution here is that offloading is shown to bite hardest on desirable difficulty: the resource that collapsed was effort on the hardest items, exactly where it is most productive. The usage-pattern analysis also complicates any simple "offloading is bad" reading. Thirty-one of the 78 codable participants barely used the tool at all, sixteen used it exploratorily — proposing answers, demanding justification, disagreeing — and they spent more time on Stage 1 than the question-copiers while both user groups scored numerically higher. As the seung-basham and misiejuk pages show for other settings, access is not the same as offloading, and interaction style may matter more than the mere presence of a chatbot.

Classroom implications and limitations

For teaching, the study does not license prohibition. It licenses attention to what a tool is doing within a task: the harm case is genAI performing the integration that is itself the learning objective, while assistance that supports exploration and evaluation looks benign or better. Two practical corollaries follow. Assess the unassisted residue, because assisted accuracy is not evidence that structure was built. And be wary of the incentive trap: promising marks or rewards for effort may not recover anything once a learner's sense of how much effort the task deserves has been anchored to tool-assisted performance — timing and Scaffolding of withdrawal, not incentives after the fact, is what the data implicate. The 82-proposition network also models the kind of material where this matters most: tasks with high element interactivity whose difficulty is inseparable from the mental model being built.

The limitations are candid and worth carrying forward. The impairment result was not consistent across levels of analysis — aggregate accuracy showed nothing, item-level modelling showed a modest effect — so it may be sensitive to analytic choices and should be replicated with larger samples or stronger incentives. The manipulation was access to the tool, not measured offloading, and with almost 40% of the genAI condition barely using ChatGPT, condition-level differences are diluted; future work should separate access from extent and type of use. Incentives moved time but never accuracy, which leaves open whether this is a weakness of the manipulation, of the task, or a ceiling effect from already-motivated participants, and completion time is only an indirect measure of effort. Question difficulty was estimated from this sample's own item-level accuracy, so the moderation finding warrants replication with independent difficulty estimates. The design targeted extrinsic motivation only, so nothing follows about interest, competence, autonomy or intrinsic motivation. Above all, this is a single laboratory session with a very short delay: Stage 2 followed Stage 1 immediately, the reasoning content was the same network rather than genuinely novel material, and Stage 3's memory test spanned both stages, giving everyone an extra independent pass over the information. Whether a semester of tool use produces the same effort signature — or whether interpolation of unassisted retrieval practice erases it, as the absence of a memory effect tentatively suggests — is not answered here. One caveat cuts against over-reading the classroom sample: participants were adults in Poznań, mostly with secondary education, which is why the page carries the higher education level rather than K-12.

Connected Concepts

  • Cognitive Offloading — the paper's central construct, and the distinction between offloading access to information and offloading its integration
  • Critical Thinking — relational reasoning as inference over a constructed structure rather than recall of propositions
  • Metacognition — metacognitive laziness, metacognitive decay, and the mismatch between perceived and actual effort requirements
  • Self-Regulated Learning — effort regulation as the self-regulatory process that the tool appeared to weaken
  • Desirable Difficulties — the finding that effort withdrawal concentrated on the hardest, most productive items
  • Generative AI — ChatGPT (GPT-5.2 Instant) available for Stage 1 only, removed before the unassisted stages
  • Conversational AI — the chat interface, and the four observed use patterns from question copying to exploratory challenge
  • Large Language Models (LLMs) — the large language model whose fluent output may inflate perceived competence and reset effort expectations
  • Motivation — motivation intensity theory and expected value of control as the competing effort accounts
  • Cognitive Psychology — working memory capacity, element interactivity, cognitive load, and the memory test
  • Transfer of Learning — whether a short-term assisted benefit transfers to subsequent independent reasoning
  • Problem Solving — Stage 1 and Stage 2 as inference problems over a network too large to hold at once

Connected Articles

Citation

Ratajczyk, D., Dymarska, A., Matłoka, A., Tomczyk, M., & Wiącek, M. (2026). Thinking with AI, reasoning without it: Cognitive offloading to generative AI weakens effort regulation. PsyArXiv (preprint, not peer reviewed).

Embed this page

Copy the code below to embed a chromeless version of this page in a learning management system or other website. The embedded view hides the site header, navigation, and footer.