Research Article
How university students work on assessment tasks with generative AI: matters of judgement
Walton, Bearman, Crawford, Tai and Boud (2025) ask how university students exercise judgment while working with generative AI on Assessment tasks. They conducted 26 interviews with undergraduates at one large Australian university, mainly using a scroll-back approach that revisits traces of students' historical GenAI interactions, then analyzed the material with a holistic definition of judgment and narrative analysis.
The paper shifts the question from whether students use GenAI to how they judge their way through it. Treating the judgment event as the unit of analysis, the authors interpret six categories of judgment events, from critically weighing GenAI's knowledge claims to copy-pasting outputs without judging them; judging GenAI, they conclude, is inseparable from judging one's own knowledge, deficits and contributions.
Synthesis: A qualitative, Multimodal AI interview study reconstructing judgment from students' own GenAI traces, showing that working with GenAI on assessment is a continuous act of judgment — about the tool's knowledge, about its limits, and, most consequentially, about what the student themselves contributed.
Key Findings
- Making judgments about knowledge when working with GenAI — the most common judgment event; students treated output as knowledge to be expanded, explained, recontextualized or reduced, splitting it into declarative "what is" and procedural "how to" content.
- Learning to judge GenAI through its limitations — students revised their theories of the tool as they hit its limits, as when Wren found ChatGPT could swap words but not restructure her sentences.
- Relying on GenAI for things they could not otherwise do — students ceded some judgment to cover a capability gap; some, like Ben, became less reliant over time as they learned grammar and essay structure from its edits.
- Adopting ideas with low levels of criticality — outputs entered assignments with minimal appraisal; Sally attributed weak marks to uncritical adoption and stopped, while Tess valued the ideas' novelty against little engagement with her disciplinary reading.
- Misjudging GenAI contributions as their own — students claimed to have altered GenAI text when the traces showed almost no change between what ChatGPT produced and what was submitted.
- Submitting GenAI work without judging it — plain copy-and-paste with no claim of learning or contribution, sometimes called "cheating" by students themselves: Samantha asked why she would guess rather than look up an answer when "everyone else around you is doing that."
Reconstructing judgment from traces
The methodological contribution is the scroll-back interview, borrowed from social media research (Robards and Lincoln 2017). Students brought historical GenAI interactions and scrolled back through them with the interviewer, who asked situated questions such as what changed here and why they stopped. Data collection was online via Zoom, with audio and screen recording. Seven interviewees did not bring historical interactions and used completed artifacts or recreated interactions instead; three students sent supplementary material afterwards and three students provided examples of actual GenAI interactions. Participants spanned Psychology, Arts, Education, Nursing, Optometry, Biomedical Science, Law, Business, Sociology and Computer Science.
From recruitment to six categories
Participants were self-identified motivated GenAI users recruited from one large Australian university: 11 students via a call to the institution's disability/accessibility unit and 15 students via subject convenors, selected purposefully for heterogeneity across discipline, mode of study, gender and years enrolled. Analysis ran in two phases: the team wrote a short holistic narrative analysis for each moment of a student working with GenAI, then clustered the narratives into thematically similar categories, focusing on the 30 of the 52 judgment events that related to assessment. A single event could be categorized more than one way, so the six categories overlap rather than partition the data.
Contradiction between what students said and did
The most striking insight is the gap between students' accounts and their artifacts. Two students who both identified as having learning disabilities described relying on GenAI to cover similar capability gaps, yet the team interpreted one as learning to become less reliant over time and the other as not noticing their reliance. GenAI can therefore enhance or hinder learning depending on circumstance, and students' own views cannot distinguish the two. Connecting this to Urban et al. (2024), whose experiment found GenAI use can raise self-efficacy while interfering with performance estimation, the authors argue the deeper risk is misjudging oneself rather than misjudging AI output.
What this means for practice
- Instructors. If tasks can be completed with GenAI in ways that circumvent the intended learning — the online quizzes described by Samantha and Rashid — copy-and-paste should not be a surprise; the authors argue the prevailing focus on cheating is more usefully reframed as a challenge for learning.
- Students. Learners need support not only to know about GenAI and how it might be used, but to understand and reflect on their relationships with knowledge, with technology, and with themselves.
- Design assessments that include evaluative judgment of the student's own contribution, not only of the GenAI contribution; interim submission steps are proposed to counter what Fan et al. (2025) call "metacognitive laziness."
- Attend to how assessment positions knowledge — whether tasks reinforce a separation between content and process acquisition, or bring the two together through explicit teaching tied to task design.
Limitations
- All 26 interviews come from one large Australian university, and recruitment was purposive and weighted toward supported students (11 students via the disability/accessibility unit), so the sample is neither large nor representative of undergraduates generally.
- The authors state the analysis was highly interpretive, blending participant accounts with researcher interpretations of their assessments and GenAI traces, with moments of unease when the two contradicted.
- Seven interviewees did not bring historical interactions, so part of the evidence rests on recreated interactions or completed texts rather than traceable GenAI history.
- Interpretation stayed contested within the team: on one student's reliance, one member argued the trade-off to learning might be offset by more consistent longer-term study, while others felt it was almost inevitably precluding learning.
Connected Concepts
- Assessment
- Academic Integrity
- Generative AI
- Higher Education
- AI Literacy
- Reducing AI Misuse
- Cognitive Offloading
- Authentic Assessment
- Student Experience
- AI Misuse and Learning Harm
Connected Articles
- Same tool, different work: patterns of generative AI use and academic outcomes — Patterns of GenAI use (evaluative integration vs low-verification uptake)
- Generative AI across the disciplines: an activity theory perspective on undergraduate students' AI use and disclosure practices — Disciplinary differences in GenAI use and disclosure
- Assessing the Quality of AI-Generated Exams: A Large-Scale Field Study — AI-generated exams and assessment quality
- To disclose or not to disclose: Peer influence and psychological factors in students' use of generative artificial intelligence — Disclosure and peer influence in GenAI use
Citation
Walton, J., Bearman, M., Crawford, N., Tai, J., & Boud, D. (2025). How university students work on assessment tasks with generative AI: matters of judgement. Assessment & Evaluation in Higher Education.