On this page

Synthesis: In a controlled between-subjects experiment, 55 undergraduate computer science students at one university completed three introductory C programming tasks with or without ChatGPT-4.5. The AI-assisted group scored 20 points higher on the coding assessment (89% vs. 69%) yet recalled less of the same material immediately (41% vs. 53%) and 48 hours later (39% vs. 52%), and attributed only 45% of the submitted code to themselves against 81% in the no-AI group. Because the loss of recall information over 48 hours did not differ between conditions, the gap reads as weaker encoding during AI-assisted work rather than faster forgetting, which puts Generative AI assistance and artifact-based Assessment in direct tension.

Key Findings

  1. Assisted performance rose while retention fell. The ChatGPT group scored 89% on the coding assessment against 69% without AI, but 41% vs. 53% on the immediate cued-recall quiz and 39% vs. 52% at 48 hours.
  2. Ownership halved. Students who worked with ChatGPT attributed 45% of the submitted code to themselves, compared with 81% in the conventional web-search condition.
  3. Mental effort rose less across tasks under AI. Self-reported effort increased less across the three tasks in the ChatGPT condition (Holm-adjusted p = .047), the load signature that Cognitive Offloading research predicts.
  4. The physiological channels were inconclusive. Confirmatory pupillometry and heart-rate-variability tests detected no significant differences in trajectories between conditions, and the authors report substantial data loss that limits their interpretation.
  5. Four of 59 participants were excluded for protocol reasons — one AI-assigned participant never used the assistant and three did not attempt the third task — leaving 29 AI-assisted and 26 no-AI students in the analysis.
  6. Forgetting rates did not differ. Both groups lost roughly two points of recall over 48 hours, so the AI group's deficit was already present at the immediate quiz: lower initial encoding rather than accelerated decay.
  7. Prior experience was covaried. Prior programming experience (any language, and C specifically) was modeled as a covariate on performance and retention, and prior ChatGPT familiarity on ownership, so the headline gaps are adjusted estimates.

Coding scores and recall move in opposite directions

The result that makes this study worth reading is not that AI helps students finish programming tasks faster; several meta-analyses already report that. It is that the two outcome measures diverged inside the same session, with the same students. Participants solved introductory C problems and then answered a ten-question cued-recall quiz on the material, first immediately and again 48 hours later. The Generative AI group's coding advantage was large, and its recall disadvantage was present at both measurement points, which rules out differential forgetting as the explanation. The design also pseudo-randomly assigned participants to the two conditions so that gender was balanced across them.

Ownership and the artifact-as-proxy problem

Students in the ChatGPT condition claimed less than half of the code they submitted (45%), against 81% without AI, a difference the authors treat as psychological ownership rather than a grading artifact. The pedagogical problem is that most programming assessment assumes the submitted artifact is a sufficient proxy for what a student knows, and this experiment is a direct test of that assumption: the artifact improved while the underlying recall fell. Instructors who grade only the working program therefore see the strongest scores from the students with the weakest retention, which is a validity problem before it is an AI problem.

What the load measures could and could not show

Cognitive load was measured on four channels: self-reported effort and difficulty (Paas scale), mean pupil diameter from a research-grade eye tracker, heart-rate variability, and task performance. Only the self-report channel produced a significant between-condition result, with effort rising less across tasks under AI assistance. The authors state plainly that the confirmatory physiological tests detected no significant trajectory differences and that substantial data loss limits their interpretation, so the load story rests on a single self-report measure — a measured claim rather than a general one about the physiology of cognitive load. The design is nonetheless notable for trying.

What this means for practice

  • Instructors. Grade learning, not only the artifact: follow AI-assisted coding work with an unaided recall or explanation task, because submitted code substantially overstates what students in this study could reproduce two days later.
  • Assessment designers. Treat ownership as a signal worth collecting. A short self-report of how much of the work the student attributes to themselves tracked the retention outcome here at no measurement cost.
  • Researchers. The 48-hour window and the 55-student single-site sample bound what this design can support; a longer retention interval and a delayed-transfer task are the obvious next studies.

Limitations

  • One university, 55 of 59 recruited students analyzed, with 29 in the AI condition and 26 without, so the study is powered for large effects only.
  • The confirmatory physiological load channels lost substantial data and detected no significant differences; the cognitive-load conclusion rests on self-reported effort.
  • Retention was measured at 48 hours only, so nothing here shows whether the recall gap persists, widens, or closes with practice.
  • Assistance was ChatGPT-4.5, and no data-collection window is reported, so the size of the performance advantage is specific to that model generation.

Connected Concepts

Connected Articles

Citation

Bergh, C., Tag, B., Vassar, A., & Renzella, J. (2026). Your Programming Students' Cognition with ChatGPT: Higher Performance, Lower Retention, and Reduced Ownership. arXiv:2609.21194.

Embed this page

Copy the code below to embed a chromeless version of this page in a learning management system or other website. The embedded view hides the site header, navigation, and footer.