π Full text: Education Sciences (MDPI, OA) Β· local
Kickbusch, Ashford-Rowe, Kemp, Boreland & Huijser (2025) argue the dominant institutional response to generative AI in assessment β surveillance and AI detection β misdiagnoses the problem: in an AI-mediated world, authenticity cannot be policed into existence; it must be redesigned. They reconceptualise authenticity as constructed where AI is expected, declared, and scrutinised, and offer discipline-agnostic "design for learning" patterns that position AI as a collaborator rather than a cheating application.
The case against detection
Detection-led responses face well-documented limits: validity and fairness failures (bias against non-native writers), notable error rates, erosion of trust, and distraction from assessment design. Detection should be a limited, situational tool β not a strategy of first resort. The constructive question is not "how do we prevent students from using AI?" but "how do we enable them to use it thoughtfully, responsibly, and effectively in contexts that mirror their future work?" Excluding AI from assessment creates an inauthentic scenario: the authentic professional justifies when, how, and why they use tools, and critically evaluates their outputs.
Authenticity as a four-dimensional continuum
Not a binary but a continuum across four intersecting dimensions:
1. Taskβcontext alignment with contemporary professional practice β judgement, decision-making, and problem-solving under uncertainty, not superficial workplace replication 2. Foregrounding professional judgement and ethics β sustainable assessment (Boud & Soler 2016), collaboration (Boud & Bearman 2024), UNESCO 2023 capability framing 3. Visibility of process β iteration, critique, rationale; polished outputs can mask superficial understanding, so assessment must reveal the "messiness" of authentic professional work 4. Appropriate use of tools (including AI) within human decision-making β tools as enablers of higher-order capability, not substitutes for it
Stage-appropriate authenticity: early units get constrained, well-scaffolded tasks; later units open complexity, uncertainty, and stakeholder engagement.
Design patterns ("design for learning moves")
- Critique, adapt, verify AI outputs: business students interrogate a chatbot-generated market analysis; pre-service teachers evaluate AI-produced lesson plans for inclusivity and pedagogical soundness; journalism students edit an AI news brief to identify bias; health students appraise AI diagnostic recommendations
- Process transparency artefacts: process logs, AI prompt records, drafts showing iterations β "behind the scenes" evidence submitted alongside the final output
- Reflective commentaries: explain key decisions, justify tool use, account for changes, with explicit criteria for depth, criticality, and ethical awareness
- Oral defences / annotated portfolios / recorded walkthroughs: probe reasoning in real time, mirroring professional practices like pitching and peer review
- Self-critique and peer feedback for feedback literacy (Boud & Molloy 2012)
- Progressive release across a programme: transparency artefacts + short defences β collaboration and negotiated briefs β capstones with external stakeholders and negotiated criteria
Challenges and institutional responsibilities
- Equity: unequal access to tools deepens divides; institutional provision (fenced AI deployments) reduces back-channel inequality; authentic formats can create new barriers (workload, carer/employment constraints) β mitigate with workload modelling, staged scaffolding, modality choice
- Ethics and bias: tools reproduce cultural stereotypes and can be fluent yet unfaithful (Bender et al. 2021); institutions should run privacy/data-protection impact assessments (PIA/DPIA) for assessment AI, vet tools against privacy/bias/accessibility criteria, and standardise prompt-log conventions that evidence process without exposing personal data
- Load and feasibility: process artefacts and defences raise workload; needs modelling and scaffolds
- Staff development: design-led collaboration rather than superficial tool training
Connections to the wiki
- The strongest critique of ai-detection in the wiki's corpus: detection cannot secure assessment-validity in an AI-mediated world
- Directly extends authentic-assessment (Zhan, Boud & Du 2025) and pairs with authentic-products-authenticated-processes-2026 (Tsiligkiris 2026 β the papers cite each other; both converge on process visibility as the validity strategy)
- "Design for learning" patterns resonate with tool-invariant-framework-agentic-ai (verification-gated oral defence) and human-in-the-loop designs
- Digital discernment connects to ai-literacy and counters over-reliance
- Reflective artefacts echo the metacognition and self-regulated-learning machinery of the wiki
- Equity/inclusion analysis connects to equity and the inclusive-design critique of care-full-feedback-genai
Related Pages
- authentic-assessment β the foundational scoping review this paper operationalises for AI
- authentic-products-authenticated-processes-2026 β Tsiligkiris (2026): six-dimension framework, authenticated processes
- ai-detection β the approach this paper argues against
- assessment-validity β why product-only evidence fails under GenAI
- ai-literacy β digital discernment as assessed capability
- academic-integrity β from policing to design
- human-in-the-loop β AI in the loop, judgement with the student
- feedback-loop β reflective commentary and feedback-use evidence
- equity β access and representational fairness
- metacognition β reflective artefacts making thinking visible
- over-reliance β uncritical tool use as the assessed failure mode
Sources
- Kickbusch, S., Ashford-Rowe, K., Kemp, A., Boreland, J., & Huijser, H. (2025). Beyond Detection: Redesigning Authentic Assessment in an AI-Mediated World. Education Sciences, 15(11), 1537. DOI