๐ Full text: AEHE (T&F, OA) ยท local
Tsiligkiris (2026) reframes authentic assessment as an evidential and validity-oriented design problem in AI-rich higher education. His central distinction โ authentic products vs. authenticated processes โ argues that assessment validity under generative AI depends not on realistic outputs alone but on architectures that make human judgement, verification, and responsibility visible. A systematic conceptual review of 37 sources yields a six-dimension framework for redesigning assessment briefs at module and programme level.
The core argument
Generative AI has not created the need for authentic assessment โ it has made weaknesses in assessment design harder to ignore. Polished products can now be generated or substantially mediated by tools, so product resemblance is an increasingly unreliable signal of capability. Tsiligkiris calls this risk construct substitution: an AI-generated or AI-mediated product is attributed to the student, producing an inference about capability that reflects the tool's performance rather than the student's learning. The evidence is concrete: fully AI-generated submissions passed through a live university examination system largely undetected (Scarfe et al., 2024), and experienced markers do not reliably detect GenAI-authored work (Kofinas, Tsay & Pike, 2025).
Crucially, the evidential question survives any AI policy: whether AI use is prohibited, permitted, or required, the assessment must still generate evidence that warrants the inference being drawn โ either about unaided capability, or about the capacity to direct, critique, and take responsibility for AI-supported work.
Six-dimension framework
1. Contextual fidelity and consequential relevance โ realistic problems, roles, constraints, audiences, artefacts; credibility to disciplinary or civic practice plus an intelligible reason to care (stakeholder, public output, policy brief), not surface imitation. 2. Cognitive demand and evaluative judgement โ analyse, evaluate, synthesise, interpret ambiguity, make defensible trade-offs. Evaluative judgement (Sadler 1989; Tai et al. 2018) matters more when machines can generate plausible first drafts. 3. Process transparency and assessment integrity โ reasoning, iteration, feedback use, and verification visible enough to support warranted inference (Boud 2000; Kane 2013). Architectures: staged submissions, annotated decision rationales, oral defence, feedback-use statements, process records. 4. Student agency and bounded choice โ topic/modality/case/medium choice within a common architecture, bounded by clear standards to preserve comparability and fairness. 5. Inclusivity and representational fairness โ realism can privilege particular communication styles, professional norms, and cultural capital; needs transparent criteria, scaffolding, equivalent routes to demonstrate achievement, and attention to whose realities are represented. 6. AI-aware validity and ethical practice โ specify how AI relates to intended outcomes, permitted uses, and the evidence supporting defensible attribution of capability; evidential requirement (does the task still capture the construct?) plus ethical requirement (disclose, critique, verify, take responsibility).
Operational tool
The framework becomes a review instrument for assessment briefs: teams examine each dimension for evidence generated, validity risks, and redesign priorities โ using a 4-point indicative alignment scale (weak โ partial โ substantial โ strong). Not every assessment must maximise all six dimensions; across a programme, tasks may emphasise different ones. It complements (rather than replaces) the AI Assessment Scale by treating AI permissions as part of a defensible assessment argument, alongside fairness, agency, cognitive demand, and process evidence.
Connections to the wiki
- Directly extends authentic-assessment (Zhan, Boud & Du 2025): the six-dimension framework retains that review's agency and collaboration insights but adds process transparency, inclusivity, and AI-aware validity as contemporary requirements
- Joins the assessment-validity critique: GenAI turns product-only evidence into an inference hazard (cf. detection limits)
- The "authenticated processes" turn aligns with human-in-the-loop designs and the verification-based defenses in tool-invariant-framework-agentic-ai and llm-fallacy-misattribution
- Process transparency resonates with formative-assessment and feedback-loop architectures that make reasoning visible
- Echoes over-reliance concerns: polished-but-superficial AI-mediated outputs as a validity threat
- Inclusion dimension connects to equity and the "matters of care" stance of care-full-feedback-genai
Related Pages
- authentic-assessment โ the non-AI scoping review (Zhan, Boud & Du 2025) this framework builds on
- beyond-detection-authentic-assessment-ai-2025 โ Kickbusch et al. (2025), the design-not-detection companion piece
- assessment-validity โ construct substitution as a validity failure mode
- ai-detection โ why detection-based responses cannot secure validity
- ai-literacy โ digital discernment as an assessed capability
- human-in-the-loop โ judgement and verification visible in the assessment architecture
- formative-assessment โ process evidence and feedback-use statements
- over-reliance โ uncritical AI-mediated production as the risk being designed against
- equity โ representational fairness dimension
- academic-integrity โ AI policy embedded in assessment architecture, not a stand-alone rule
- feedback-loop โ sustainable evaluative judgement
Sources
- Tsiligkiris, V. (2026). From authentic products to authenticated processes: a systematic conceptual review of authentic assessment in AI-rich higher education. Assessment & Evaluation in Higher Education. DOI