Concept
Authentic Assessment
Authentic assessment — the design of assessments that examine student performance on worthy, realistic intellectual tasks, rather than isolated, standardized test items. Originating with Wiggins (1990) as a counterbalance to standardized tests, authentic assessment has evolved from replicating workplace tasks toward a multi-dimensional framework encompassing professional, digital, personal, and social authenticity. Generative AI has made authentic assessment newly essential: any task a language model can credibly simulate in a take-home setting loses its validity as evidence of original student competence, so authentic forms must be redesigned around what AI cannot credibly counterfeit.
Questions to Consider
- Authentic assessment examines performance on worthy, realistic intellectual tasks rather than isolated test items. Before reading, did you assume an 'authentic' task simply replicates a real-world or workplace scenario? This page argues authenticity has evolved well beyond mere replication — what other forms can it take?
- A central claim is that any task a language model can credibly simulate in a take-home setting loses its validity as evidence of original student competence. Which assessment in your own teaching or study would you now consider 'counterfeitable' — and what makes you say that?
- The page highlights a striking gap: only 3 of 37 studies addressed social authenticity — whether assessment helps students contribute to societal transformation. Why do you think the social dimension of authenticity is so neglected, and what might assessing it actually look like?
- The knowledge base argues authenticity must be 'redesigned, not policed.' When students co-author rubrics, choose what and how to submit, and engage in real-time interaction, AI use becomes expected and declared rather than concealed. How would that shift the relationship between student and assessor?
- Oral exams and process-based portfolios are offered as AI-resistant authentic forms. But graded reflection and self-assessment risk becoming performative — students performing the reflection they think is wanted. How could you design for genuine reflection rather than performance?
- When AI supplies the rubric, the feedback, and the monitoring, the page warns that a student's metacognitive practice may be displaced. If the tool does the evaluating, what is left for the learner to actually practice or learn?
Introduction
Authentic assessment sits at the heart of how Assessment is being rethought in the AI era. It connects to Assessment Validity (does the assessment measure what it claims?), Formative Assessment (authentic tasks that inform learning), and Academic Integrity (moving from detection to designing tasks where AI use is expected and declared). It is a central response in the knowledge base's assessment-redesign literature.
The evolution of authenticity
- 1990s origins — worthy intellectual tasks: Wiggins (1990) proposed authentic assessment as direct examination of "student performance on worthy intellectual tasks," a counterbalance to standardized tests.
- Late-1990s uptake — workplace replication: Joughin (1998) framed authenticity as the extent to which assessment replicates professional practice or real life — a view that dominated for two decades.
- 2020s critique — beyond replication: McArthur (2023) argued authentic assessment must enable students to "influence the future and transform society" rather than merely replicate existing tasks; Ajjawi et al. (2024) broadened authenticity to contextual, task, and personal forms that reflect student experience.
- Generative AI as existential challenge: Zhan, Boud & Du (2025) note that generative AI makes workplace-replication authenticity newly vulnerable, driving a pivot toward digital literacy, real-time collaboration, social contribution, and individual meaning-making that AI cannot credibly counterfeit.
A six-dimensional design model
Zhan, Boud & Du's (2025) scoping review of 37 empirical studies (2000–2024) proposes six design dimensions: (1) authenticity in assessment (assessment, professional, digital, self, and social authenticity — with only 3/37 studies addressing social authenticity, a critical gap), (2) cognitive challenges, (3) assessment criteria (with students often passive recipients rather than co-authors of rubrics), (4) feedback (formative-dominated, but sustainable feedback rare), (5) student agency (choice in what/how/when/where to submit was rare), and (6) social collaboration. It also proposes a cyclical co-design model — negotiate goals, create context, co-design criteria, plan feedback — that AI tools could operationalize.
Dollinger and Nieminen (2026) bound this optimism with the paradox of inclusive assessment: both the accommodations model and structural critiques that reframe assessment as disabling leave intact the zero-sum logic that "for one student to succeed, someone else needs to fail," so distributed and agentic forms must change that logic rather than only widen access.
Authentic assessment in the AI era
The knowledge base's assessment-redesign literature argues that authenticity must be redesigned, not policed:
-
Beyond Detection contends that authenticity cannot be policed into existence; it must be designed, positioning AI as a declared collaborator rather than a cheating application, and prioritizing authentic, process-based assessment over surveillance.
-
Responsible Assessment reframes assessment around validity evidence and authentic tasks that mirror students' future work.
-
Authentic products, authenticated processes examines how AI-rich higher education can assess both genuine outputs and the processes that produced them.
-
An AI drafting partner still leaves the judgment. Paula et al. (2026) piloted a GPT-4.1 assessment designer with eight coordinators: it produced usable overviews, tasks, criteria, timelines and rubrics but repeatedly missed disciplinary context and topic sequencing, and all eight refused end-to-end automation, locating academic judgment rather than the model.
-
The tool-invariant framework argues for assessing computational methods and process rather than tool-specific outputs, using oral defense and verification.
-
Institutional infrastructure for process evidence. A commissioned quality-assurance framework turns the process argument into policy, treating learning-process traces — time-stamped edits, resource use, metacognitive prompts — as the durable evidence a transcript cannot show, and citing process features that explain more variance in performance than product features alone (Lodge et al. (2026)).
-
Scalable oral assessment, with a scoring caveat. Pentland, Lowenthal & Krier (2026) delivered time-limited, non-revisitable recorded responses graded against embedded rubrics, and found students outscored their in-person multiple-choice exams; they stress this is a format effect rather than a learning gain, with LLM re-scoring agreeing with the instructor at ICC = 0.73 and 0.60.
-
Reconsidering oral exams positions the oral exam/assessment as a low-tech authentic alternative that is inherently AI-resistant — its real-time, interactive dialogue tests comprehension, critical thinking, and reasoning (not memorization), mirrors professional practice, and prevents students from using AI to generate and memorize answers. It offers a concrete set of practical recommendations (clear rubrics, standardized content, assessor training, prompting guidelines, bias mitigation) for reintroducing Oral Assessment across high school and higher education.
-
E-portfolio assessment is another authentic, process-based form that resists AI fabrication: Zhan, Boud & Du (2025) identify social contribution portfolios among the authentic forms most robust to generative AI, and Beyond Detection recommends annotated portfolios and recorded walkthroughs that probe reasoning in real time. Ni & Lam (2026) and Laksana et al. (2026) show generative AI can assist the portfolio process — feedback, drafting, reflection — while the portfolio's reasoning traces and drafts preserve authenticity.
-
Authenticity criteria can themselves restrain AI use. Chen and Zou (2026) found that a rubric requiring students to ground a group presentation in their own first-hand teaching experience, and to reflect on shared classroom observations, led seven of fifteen student groups to deliberately reduce their GenAI use. Students argued the tool could not meet the epistemic demand — "AI only knows that moment when you type" — because it lacked the longitudinal, situated knowledge their classmates and teacher had. Both layers mattered: the authenticity of the task and the relational authenticity of contributing one's own thinking to a group, which reframed heavy AI use as free-riding on peers. The design implication is that authentic, experience-grounded criteria do evaluative work even without enforcement — they supply a reason for restraint that policy statements cannot.
-
Twin a vulnerable task with a less-vulnerable one. Roe, Perkins & Giray (2026) pair a GenAI-vulnerable task with a second assessing the same outcomes, schedule them close together so each can cross-check the other, and make the mark interdependent through a confirmatory threshold or weighting.
Sharma (2026) argues that integrity-oriented design and authentic design are not the same thing, and that the difference is epistemic rather than stylistic. Authenticity asks whether a task mirrors worthwhile real-world practice; integrity-oriented design asks whether learners can justify their decisions and assume responsibility in relation to disciplinary standards, which makes integrity an explicit, assessable criterion embedded in the task architecture rather than an incidental by-product of realism. The practices he offers — annotated decision trails, verification of GenAI-contributed claims, oral defense and dialogic accountability, draft differences with version history — are recognizably authentic-assessment forms, but they are selected for the judgment they make visible rather than for their realism, which is a useful corrective for tasks that look authentic and remain counterfeitable in substance. His caution belongs on this page too: requiring documented reasoning privileges learners more fluent in reflective discourse, and judgment as evidence "remains relational and situated rather than mechanically verifiable", so the design carries its own interpretive-reliability load.
The same question is reframed by Thapa and Lewis (2026), who name what a task should authenticate rather than whether it looks real: epistemic authenticity, the extent to which assessment captures genuine engagement with interpretation, evaluative judgment, reasoning and knowledge construction. Their argument is that authenticity alone does not survive generative AI, because an authentic task whose evidence is one unsupervised product is exposed to the same substitution problem as the essay it replaced. Making reasoning visible instead requires staged submissions, reflective justification, dialogic engagement and evaluative transparency. They carry the caution above into design: reflective and dialogic tasks can privilege students confident in academic self-articulation unless they are inclusively designed and scaffolded, and the relational work intensifies educator emotional labor in large or resource-constrained settings.
When practitioners retreat to the policed formats
The knowledge base argues that authenticity must be designed rather than policed, and Goldstein, Marae-Haj and Zidan (2026) supply the counter-evidence from practice: after a trust crisis in which pre-service teachers submitted raw AI-generated work as their own, teacher educators in seven Israeli colleges redesigned assessment around process evidence (prompts, document version history, monitored group contributions) and in-class performance — but several also described falling back, as a last resort, on supervised examinations and anticipated oral defenses on theses, one participant calling the return to exams personally painful yet unavoidable. Read against the design-first position, the finding marks a failure mode rather than a solution: where authenticity is not redesigned in advance, the practical response to unassessable submitted work is reinstated control — surveillance by another name — which reintroduces the very formats the authentic-assessment literature tries to move beyond. It also shows the felt cost of the alternative: participants reported their take-home written assessments could no longer evidence learning at all, and that process documentation took real work to establish.
A second failure mode sits inside redesign itself: across five online language-teaching methods, authenticity-related tensions dominated teachers' accounts (11 of 20), including task realism corrosion — activities redesigned to be AI-resistant rather than authentic (Baoyi and Khan (2026)).
The retreat carries an evidentiary cost that the case-file evidence makes visible. Munoz et al. (2026) coded 1,162 generative-AI misconduct allegations and found that the strongest evidence types are generated by the investigation or by supervision rather than by the allegation, and that process evidence such as drafts, supervision meetings and presentations appears only where those practices already exist — so in an unredesigned assessment the process evidence this literature recommends is simply not available, and cases rest on weaker categories. What stands in for it is the weakest material in their corpus: detector output drew the lowest probative ratings of any category, and Hadra et al. (2026) show why, with macro accuracy of 0.69 (Originality) and 0.61 (Turnitin) across 192 texts, near-total failure on hybrid human–AI writing, and a borderline tendency to misread EFL student work as AI. Wright (2026) adds the rule-design dimension: prohibitions written by platform identity rather than function are over-inclusive by definitional accident, so a policed format can sanction a student who did not do the thing the rule was designed to prevent, falling hardest on the students who relied on transcription tools for Accessibility. Read together, the retreat substitutes the least probative evidence available for the process evidence that authentic design would have generated in the first place.
- Self-regulated learning: student agency in authentic assessment (choice, self-reflection, co-design) mirrors the forethought → performance → self-reflection cycle, though graded reflection risks becoming performative.
- Metacognition: Metacognition is required for students to evaluate their work against co-designed rubrics; when AI supplies the rubric, feedback, and monitoring, the student's metacognitive practice may be displaced.
- Formative and sustainable feedback: authentic assessment emphasizes formative, future-oriented feedback that transfers to later contexts.
Implications for AI in education
-
AI-proof assessment types: in-vivo demonstrations, social-contribution portfolios, co-created artifacts with auditable provenance, and real-time embodied interaction are more resilient to generative AI than take-home essays or MCQs.
-
The exposure is measurable. A three-year psychology degree proved 90% passable: ChatGPT produced adequate output on 36 of 40 coursework assessments, and only the four tasks requiring presence, a visual artifact, or the student's own data resisted (Ivory et al. (2026)).
-
Co-design at scale: AI tools could enable rubric co-design and student co-creation of assessment parameters at classroom or MOOC scale — though machine-mediated agency must be designed carefully.
-
Address the social-authenticity gap: only 3/37 studies addressed social issues; AI assessment tools should help students contribute to societal transformation, not merely simulate it.
-
Sustainable feedback: AI feedback should be designed to transfer to future contexts, not just provide reactive, momentary corrections.
-
Process-focused measurement now has validity evidence. Oliveira et al. (2025) scored 70 graded essays on how students steered a GenAI dialogue and made course knowledge visible; those process scores correlated r = 0.54 with traditional essay scores while rewarding conceptual work over structured task specification.
-
Authentic assessment suits practice-oriented fields. Mesny, Roberge-Maltais & Galy (2026) find authentic assessment especially well-suited to management education: tasks mirroring real professional problems (live consulting projects, dashboards with executive briefings) align with the field's practice-oriented, employability focus and can support inclusive, integrity-preserving alternatives to exam-centered assessment in the generative AI era. In their review of 58 articles from four management-education journals, however, authentic assessment appeared mainly via technology-mediated simulations and was often conflated with experiential learning — a terminology gap that can obscure its broader value and uptake.
-
Authenticity is not security, and the gap is visible at sector scale. Where the single-degree audit above found 90% of coursework passable, Villanueva (2026) scores recorded presentations, remote group projects and unsupervised digital artifacts in the high-exposure band across 53,915 items, because the public record shows nothing about how production was supervised.
Connected Concepts
- Learner Identity — evolving disciplinary, professional, creative, and academic learner identities
- E-Portfolio
- Problem-Based Learning
- Assessment
- Assessment Validity
- Formative Assessment
- Automated Assessment
- Academic Integrity
- AI Detection
- Self-Regulated Learning
- Metacognition
- AI Ed Evaluation
- Higher Education
- Generative AI
- AI in Education
- Feedback
- Summative Assessment — Summative assessment: AI-resistant formats (oral, proctored, closed-book exams)
- Arts, Design and Media Education
Connected Articles
-
Preserving epistemic authenticity: process-oriented assessment in the age of generative AI — process-oriented assessment and epistemic authenticity as what a task should authenticate
-
The Integrity of Psychology Assessments in the AI Age: A Critical Examination — What AI could not pass: presence, visual artifacts, and the student's own data (Ivory et al. 2026)
-
The Paternalistic Filter: Epistemic Injustice and Differential Refusal in LLM-Mediated History Education for Marginalized Romanian Students — Paternalistic AI use and student identity in history education
-
From Classroom Design to Newsroom Practice: Assessment Intervention Designing GenAI
-
Reimagining Success and Failure: Equitable Assessment Practices in an Age of Artificial Intelligence
-
Students' Perceptions of Multiliteracies Development Using AI-Assisted Portfolio Assessment
-
Problem-Based Learning and the Structural Conditions for Productive AI Integration
-
Designing for Authentic Assessment: A Scoping Review — Designing for Authentic Assessment: A Scoping Review
-
Beyond Detection: Redesigning Authentic Assessment in an AI-Mediated World — Beyond Detection: Authentic Assessment
-
Responsible Assessment in the AI Era: Key Insights from a Future-Focused Conference — Responsible Assessment in the AI Era
-
From authentic products to authenticated processes: a systematic conceptual review of authentic assessment in AI-rich — Authentic Products, Authenticated Processes
-
A Tool-Invariant Framework for Teaching and Assessing Computational Methods in the Age of Agentic AI — Tool-Invariant Framework for Teaching and Assessing Computational Methods
-
Confidence Estimation in Automatic Short Answer Grading with LLMs — Confidence-aware automated short-answer grading
-
The LLM Fallacy: Misattribution in AI-Assisted Cognitive Workflows — The LLM Fallacy and Misattribution of Competence
-
The University AI Didn''t Replace: Rethinking Universities in the AI Era — Rethinking Universities in the AI Era
-
Reassessing Academic Integrity in the Age of AI: A Systematic Literature Review on AI and Academic Integrity — Multiple assessment methods to counter AI misconduct
-
Reconsidering the Use of Oral Exams and Assessments: An Old Way to Move Into a New Future — Reconsidering oral exams as authentic, AI-resistant assessment
-
Assessment twins: An approach for strengthening assessment validity in the age of generative AI — Assessment twins for strengthening assessment validity in the age of GenAI (Roe, Perkins & Giray 2026)
-
Assuring quality learning in a gen AI-integrated future: The role of adaptive capabilities. TEQSA, June 2026 — Adaptive capabilities for assuring quality learning in a gen AI-integrated future (Lodge et al. 2026)
-
Asynchronous Oral Assessments: Enhancing Integrity, Engagement, and Communication in the AI Era — Asynchronous Oral Assessments in the AI Era (Pentland 2026)
-
Assessing students' DRIVE: A framework to evaluate learning through interactions with generative AI — DRIVE: assessing learning through GenAI interaction (DRI + Visible Expertise)
-
Students' Agency in GenAI-Mediated Group Assessment: An Ecological-Emergent Perspective — Authenticity criteria that led student groups to reduce GenAI use
-
Pioneering Teacher Educators Navigating AI Integration in Pre-Service Teacher Preparation: Strategies and Challenges — The AI trust crisis that pushed teacher educators back toward supervised exams and oral defenses (Goldstein et al. 2026)
-
Educational integrity in GenAI-augmented assessment: making judgment visible — Integrity-oriented design distinguished from authentic assessment: making judgment visible (Sharma 2026)
-
How strong is the evidence in generative AI-related academic misconduct allegations? A mixed-methods analysis — What misconduct allegation files actually contain as evidence, and the process evidence they lack (Munoz et al. 2026)
-
Evaluating the accuracy and reliability of AI content detectors in academic contexts — Detector accuracy and hybrid-writing failure: why detection is weak fallback evidence (Hadra et al. 2026)
-
Transcription is not generation: Distinguishing non-generative AI tool use from academic misconduct in higher education assessment — Over-inclusive AI prohibitions: transcription is not generation (Wright 2026)
-
Designing Authentic Assessments with Generative AI: A Pilot Study of Assessment Authentifire in Higher Education — Designing Authentic Assessments with Generative AI: A Pilot Study of Assessment Authentifire in Higher Education
-
AI and higher education: teachers' perceptions of students' ethical tensions in online pedagogical methods — Task realism corrosion: AI-resistant redesign eroding authenticity in online language teaching
-
How AI-vulnerable is Australian higher education assessment? A computational audit of Group of Eight Arts and Humanities units, 2022–2026 — Authentic-surfaced formats (recorded presentations, remote group projects, digital artifacts) score high across 53,915 items: authenticity does not verify authorship