Research Article
Beyond Detection: Redesigning Authentic Assessment in an AI-Mediated World
Synthesis: Kickbusch, Ashford-Rowe, Kemp, Boreland, and Huijser (2025) argue the dominant institutional response to generative AI in Assessment — surveillance and AI detection — misdiagnoses the problem: in an AI-mediated world, authenticity cannot be policed into existence; it must be redesigned. They reconceptualize authenticity as constructed where AI is expected, declared, and scrutinized, and offer discipline-agnostic "design for learning" patterns that position AI as a collaborator rather than a cheating application.
The case against detection
Detection-led responses face well-documented limits: validity and fairness failures (bias against non-native writers), notable error rates, erosion of trust, and distraction from assessment design. Detection should be a limited, situational tool — not a strategy of first resort. The constructive question is not "how do we prevent students from using AI?" but "how do we enable them to use it thoughtfully, responsibly, and effectively in contexts that mirror their future work?" Excluding AI from assessment creates an inauthentic scenario: the authentic professional justifies when, how, and why they use tools, and critically evaluates their outputs.
Authenticity as a four-dimensional continuum
Not a binary but a continuum across four intersecting dimensions:
- Task–context alignment with contemporary professional practice — judgment, decision-making, and Problem Solving under uncertainty, not superficial workplace replication
- Foregrounding professional judgment and Ethics — sustainable assessment (Boud & Soler 2016), collaboration (Boud & Bearman 2024), UNESCO 2023 capability framing
- Visibility of process — iteration, critique, rationale; polished outputs can mask superficial understanding, so assessment must reveal the "messiness" of authentic professional work
- Appropriate use of tools (including AI) within human decision-making — tools as enablers of higher-order capability, not substitutes for it
Stage-appropriate authenticity: early units get constrained, well-scaffolded tasks; later units open complexity, uncertainty, and stakeholder engagement.
Design patterns ("design for learning moves")
- Critique, adapt, verify AI outputs: business students interrogate a chatbot-generated market analysis; pre-service teachers evaluate AI-produced lesson plans for inclusivity and pedagogical soundness; journalism students edit an AI news brief to identify bias; health students appraise AI diagnostic recommendations
- Process transparency artifacts: process logs, AI prompt records, drafts showing iterations — "behind the scenes" evidence submitted alongside the final output
- Reflective commentaries: explain key decisions, justify tool use, account for changes, with explicit criteria for depth, criticality, and ethical awareness
- Oral defenses / annotated portfolios / recorded walkthroughs: probe reasoning in real time, mirroring professional practices like pitching and peer assessment
- Self-critique and peer feedback for feedback literacy (Boud & Molloy 2012)
- Progressive release across a program: transparency artifacts + short defenses → collaboration and negotiated briefs → capstones with external stakeholders and negotiated criteria
Challenges and institutional responsibilities
- Equity: unequal access to tools deepens divides; institutional provision (fenced AI deployments) reduces back-channel inequality; authentic formats can create new barriers (workload, carer/employment constraints) — mitigate with workload modeling, staged scaffolding, modality choice
- Ethics and bias: tools reproduce cultural stereotypes and can be fluent yet unfaithful (Bender et al. 2021); institutions should run privacy/data-protection impact assessments (PIA/DPIA) for assessment AI, vet tools against Privacy/bias/Accessibility criteria, and standardize prompt-log conventions that evidence process without exposing personal data
- Load and feasibility: process artifacts and defenses raise workload; needs modeling and scaffolds
- Staff development: design-led collaboration rather than superficial tool training
What this means for practice
- Instructors. Replace detection-first rules with tasks where tool use is expected and declared: have students critique, adapt, and verify AI outputs (a chatbot-generated market analysis; an AI-produced lesson plan) and grade the judgment they show, not the artifact they submit.
- Instructors. Build visibility of process into the assessment itself — prompt logs, draft iterations, short reflective commentaries with published criteria for depth, criticality, and ethical awareness — so a polished output cannot stand in for understanding.
- Instructors. Stage authenticity deliberately: constrained, scaffolded tasks in early units, then open complexity, uncertainty, and stakeholder engagement later, with equivalent-standards modality choice for students who need it.
- Administrators. Stop treating detection as a strategy of first resort: fund redesign, run privacy/data-protection impact assessments on assessment AI, and vet tools against Privacy, bias, Accessibility, and auditability criteria.
- Administrators. Provide fenced institutional AI access and model the workload before adding process artifacts and oral defenses, so authentic formats do not become new barriers for students with carer or employment constraints and do not deepen equity gaps.
Limitations
- This is a conceptual/position paper (theoretical analysis) whose two contributions are a reconceptualization of authenticity and a set of discipline-agnostic design patterns; it reports no participants, no course, and no trial, so it supports design reasoning rather than evidence that the patterns improve learning.
- The four-dimensional authenticity continuum and the design moves are derived from prior literature and the authors' own practice in a single institutional learning-and-teaching unit, not from any measurement of authenticity or of student outcomes.
- The paper's own challenges section concedes the proposed formats raise workload and can create new barriers (carer and employment constraints, real-world simulation limits) without offering cost or feasibility data; detection's error rates and bias are likewise cited from other studies rather than re-tested here.
Citation
Kickbusch, S., Ashford-Rowe, K., Kemp, A., Boreland, J., & Huijser, H. (2025). Beyond Detection: Redesigning Authentic Assessment in an AI-Mediated World. Education Sciences, 15(11), 1537. DOI 10.3390/educsci15111537.