On this page

Synthesis: Hutchison et al. (2026) develop and validate a method for measuring critical engagement with AI code completion tools in educational settings. Using behavioral signals (time-to-accept, edit distance from suggestion) and embedded attention checks, they find that the majority of students accept AI code suggestions passively, without critically evaluating correctness or appropriateness. This 'tab-and-go' behavior directly threatens the development of programming skills, as students bypass the cognitive effort required for Transfer of Learning. The work provides a methodological toolkit for Formative Assessment in AI-augmented programming courses, enabling instructors to detect when students are over-reliant on AI. The findings connect to Student Experience research in STEM Education by showing that the mere availability of AI tools does not lead to productive learning — structured pedagogical interventions are required to ensure students engage critically rather than deferring to AI output.

What this means for practice

  • Instructors. Judge critical engagement from behavior, not from whether work looks finished. Students in this study tab-accepted an average of 17.1 suggestions, and only 22 of 55 passed all 26 test cases — a completed solution is not evidence that the suggestion was evaluated.
  • Instructors. Read dwell time and accept-then-modify rates as early warning signals. Average dwell time before first interaction with a suggestion was 12.2 seconds, and code execution was the strongest single correlate of task performance (ρ = 0.26), so students who never run or revise what they accepted are the ones to check on.
  • Instructors. Use deliberately incorrect suggestions as a teaching device rather than only as an assessment trap. Clover's injected errors functioned as an attention check in this study, but they can equally be turned into a Formative Assessment moment with immediate feedback.
  • Designers. Surface reliance signals in the tool itself — dwell time, edits after acceptance, deletions — so students can see their own offloading patterns while they can still change them.
  • Instructors. Set explicit norms for the lab before students open the tool, since the course normally prohibited generative AI and the exception itself may have shaped how students treated the suggestions.

Limitations

  • 56 students consented across four lab sections of one Java course at a single university; one was dropped after a dwell time more than three standard deviations from the mean, leaving 55 in the analysis.
  • The task was a single 60-minute session on the rainfall problem, and task performance had limited variability: some students submitted non-compiling solutions while others passed all 26 test cases, with little in between; the system also used one AI model in one tool (Clover), though model choice affects both suggestion quality and latency.
  • Attention checks injected deterministic incorrect suggestions, which the authors note do not replicate the spontaneous hallucinations students face in real coding work, and suggestion type or complexity was not recorded.
  • Interaction logging stops at the first action after a suggestion appears, so engagement that continues after acceptance (or deletion) was not measured, and participation credit did not depend on performance, so the authors cannot rule out students using external tools such as Google or ChatGPT.

Citation

Jessica Hutchison, Ian Tyler Applebaum, Kenneth Angelikas, Kush Rakesh Patel, Phuoc Nguyen, Antonio Lazaro, Nicholas Rucinski, Rahad Arman Nabid, Stephen MacNeil (2026). To Tab or Not to Tab: Measuring Critical Engagement in AI Code Completion Tools Using Behavioral Signals and Attention Checks. cs.HC (ITiCSE 2026).

Embed this page

Copy the code below to embed a chromeless version of this page in a learning management system or other website. The embedded view hides the site header, navigation, and footer.