Research Article
How university students work on assessment tasks with generative AI: matters of judgement
How university students work on assessment tasks with generative AI: matters of judgement — a qualitative Multimodal study of 26 Australian university students using a scroll-back interview approach to reconstruct how students exercise judgement as they work with GenAI on assessment tasks. Walton, Bearman, Crawford, Tai & Boud (2025) identify six categories of judgement events, revealing a wide spectrum — from critically evaluating GenAI knowledge to uncritically submitting AI content.
Walton et al. (2025) shift the assessment debate from whether students use GenAI to how they judge their way through working with it. Using the scroll-back method (revisiting students' actual historical GenAI interactions during interviews), the study offers a fine-grained account of student decision-making in real assessment contexts — spanning responsible, critical use and problematic shortcuts.
Method
- Design: Qualitative, multimodal; scroll-back interviews that revisit traces of students' historical interactions with GenAI.
- Sample: 26 Australian university students.
- Analysis: Holistic definition of judgement; narrative approach.
Key Findings — six categories of judgement events
- Making judgements about knowledge when working with GenAI — evaluating the accuracy and trustworthiness of AI output.
- Learning to judge GenAI through its limitations — calibrating expectations from AI's errors and gaps.
- Relying on GenAI for things they could not otherwise do — using AI to extend capacity beyond current skill.
- Adopting ideas with low levels of criticality — accepting AI output without scrutiny.
- Misjudging GenAI contributions as their own — failing to distinguish AI-generated from self-generated work.
- Submitting GenAI content in an assignment — in some cases directly submitting AI output.
The categories span a continuum from productive judgement (1–3) to problematic or unreflective use (4–6), showing that student GenAI use in assessment is heterogeneous and shaped by individual judgement capacity.
Implications
- For Assessment and Academic Integrity: students exercise a range of judgements, not a binary of compliant vs. dishonest use — integrity frameworks must recognize this spectrum.
- For AI Literacy and Reducing AI Misuse: the findings point to judgement (evaluating AI knowledge, calibrating trust, recognizing AI's contributions) as a teachable capacity, connecting to the performance–learning gap and over-reliance literature.
- For Authentic Assessment: because AI is now embedded in everyday assessment practice, redesigning tasks to require and reward judgement is central.
Connected Concepts
- Assessment
- Academic Integrity
- Generative AI
- Higher Ed
- AI Literacy
- Reducing AI Misuse
- Cognitive Offloading
- Authentic Assessment
- Student Experience
- AI Misuse Learning Harm
Connected Articles
- Stamatoulis GenAI Use Patterns 2026 — Patterns of GenAI use (evaluative integration vs low-verification uptake)
- Jiang GenAI Activity Theory Disciplines 2026 — Disciplinary differences in GenAI use and disclosure
- Assessing Quality AI Generated Exams Field 2025 — AI-generated exams and assessment quality
- Qu Wang Disclose Or Not GenAI 2026 — Disclosure and peer influence in GenAI use
Citation
Walton, J., Bearman, M., Crawford, N., Tai, J., & Boud, D. (2025). How university students work on assessment tasks with generative AI: matters of judgement. Assessment & Evaluation in Higher Education.