---
source_url: https://doi.org/10.1080/02602938.2026.2695376
ingested: 2026-08-03
sha256: 988f98c3e6c3f32fa9f7e99395028b009f2b67989ee45fae3babcf190b4a121b
---

# From authentic products to authenticated processes: a systematic conceptual review of authentic assessment in AI-rich higher education

Vangelis Tsiligkiris (Nottingham Business School, Nottingham Trent University). *Assessment & Evaluation in Higher Education* (Taylor & Francis), published online 04 Jul 2026. DOI: 10.1080/02602938.2026.2695376. Open access (CC BY).

## Abstract (summary)

Systematic conceptual review of authentic assessment in higher education developing a six-dimension framework for assessment design in digitally mediated and AI-rich conditions. Retained corpus of 37 substantive sources. The synthesis shows authentic assessment should not be reduced to workplace simulation nor treated primarily as a response to academic misconduct; it is better understood as a multidimensional design orientation. Main contribution: distinguishing **authentic products from authenticated processes** — assessment validity under generative AI depends not only on realistic outputs but on architectures that make human judgement, verification, and responsibility visible.

## Method

- Systematic conceptual review (Snyder 2019; Torraco 2005), for construct clarification/framework development, not effect-size aggregation
- Sources: Scopus, Web of Science Core Collection, ERIC (searched January 2026), plus backward/forward citation chaining and targeted supplementary searching (validity theory, process evidence, GenAI redesign)
- Eligibility: higher education, substantive contribution to definition/design of authentic assessment; n = 37 retained
- Coding framework: authenticity definitions/characteristics (15), task design/alignment/performance architecture (14), others incl. inclusion, evaluative judgement, AI-mediated validity
- Single author; no independent screening/coding or inter-rater reliability; Anglophone-weighted, concept/review-heavy corpus

## The validity challenge

- Classical critique: conventional assessment measures what is easiest to examine rather than most educationally valuable (Wiggins 1998; Ashford-Rowe et al. 2014)
- GenAI has not created the need for authentic assessment but made design weaknesses harder to ignore; prohibition/disclosure/detection responses questioned (Cotton, Cotton & Shipway 2024; Luo 2024)
- Growing position: GenAI is a legitimate component of professional practice; assessment should evaluate competent AI-supported performance (Fawns et al. 2025); Corbin et al. (2026) call AI+assessment a "wicked problem" requiring situated judgement under permanent uncertainty
- **Construct substitution**: risk that an AI-generated/mediated product is attributed to the student → inference about capability reflects the tool, not the learning (Scarfe et al. 2024: AI submissions passed live exams undetected; Kofinas, Tsay & Pike 2025: experienced markers cannot reliably detect GenAI work)
- Whether AI is prohibited, permitted, or required, assessment must generate evidence warranting the inference (about unaided capability OR capacity to direct/critique/take responsibility for AI-supported work)

## Six-dimension framework

1. **Contextual fidelity and consequential relevance** — realistic problems, roles, constraints, audiences, artefacts; credibility to disciplinary/civic practice + an intelligible reason to care (external stakeholder, public-facing output, policy recommendation, design brief); not surface imitation
2. **Cognitive demand and evaluative judgement** — analyse, evaluate, synthesise, interpret ambiguity, defensible trade-offs; poorly designed realistic tasks can still be cognitively thin; links to evaluative judgement (Sadler 1989; Tai et al. 2018), more important when machines generate plausible first drafts
3. **Process transparency and assessment integrity** — reasoning, iteration, feedback use, decision-making, verification visible enough to support warranted inference (Boud 2000; Kane 2013); transparent architectures: staged submissions, annotated decision rationales, oral defence, feedback-use statements, process records, context-specific questioning
4. **Student agency and bounded choice** — topic/modality/case/audience/medium/data-source choice within a common architecture; bounded by clear standards to preserve comparability and fairness
5. **Inclusivity and representational fairness** — authenticity is not an unqualified good; realistic tasks may privilege particular communication styles, professional norms, cultural capital, prior access (Forsyth & Evans 2019; Tai et al. 2023); need transparent criteria, scaffolding, equivalent routes to demonstrate achievement, conscious attention to whose realities are represented
6. **AI-aware validity and ethical practice** — specify how AI tools relate to intended outcomes, permitted/prohibited uses, and what evidence supports defensible attribution of capability; evidential requirement (does the task still capture the construct when AI can generate/transform/refine work) + ethical requirement (disclose, critique, verify, take responsibility); not all-AI-prohibited, not all-oral/invigilated/process-heavy; deliberate design for human judgement in relation to machine assistance

## Relationship to prior frameworks (Table 3)

Retains Wiggins (1998), Gulikers et al. (2004), Ashford-Rowe et al. (2014), Villarroel et al. (2018) insights (task fidelity, challenge, performance, transfer) while foregrounding process transparency, inclusivity, and AI-aware validity as contemporary requirements. Contribution is synthesis, not individual novelty.

## Operationalising (review tool)

- Module/programme team reviews each brief against the six dimensions: evidence generated, validity risks addressed, redesign priorities
- Not every assessment must maximise all six dimensions; across a programme tasks may emphasise different dimensions (Boud & Bearman 2024); choices deliberate, proportionate, defensible
- **Indicative alignment scale (Table 5)**: 1 weak → 2 partial → 3 substantial → 4 strong alignment; heuristic for deliberation, not a psychometric instrument
- Illustrative application: policy-briefing assessment (recommend evidence-based intervention to a local authority) — strong fidelity/cognitive demand if students evaluate competing options and defend recommendations; process transparency via staged problem-scoping statement, annotated evidence review, oral defence; may require redesign on inclusion (policy genre familiarity) and AI-aware validity (specify AI use, require disclosure/verification/responsibility)
- Complements (not replaces) the AI Assessment Scale: places AI-aware design alongside fairness, agency, cognitive demand, process evidence; AI permissions as part of a defensible assessment argument, not a stand-alone policy choice (Ajjawi et al. 2024; Kickbusch et al. 2025; Perkins et al. 2024; Su et al. 2026)

## Discussion

- Reframes authenticity as an evidential and validity-oriented design problem: what evidence does the assessment generate, what capability can it support, what risks weaken the inference
- Key shift: from "does a task look authentic" to "what kind of authenticity does it enact, for whom, with what evidence, with what consequences"
- Authentic assessment is not valuable because it defeats AI or prevents misconduct, but because it creates conditions for students to demonstrate situated judgement and take responsibility in ways supporting warranted inference
- Gaps: evidence base conceptually expansive but operationally inconsistent; limited cross-disciplinary empirical evaluation of redesigned authentic assessments with shared outcome measures
- Use critically: process evidence improves interpretability but raises workload; high-fidelity tasks may reproduce unequal access unless deliberately scaffolded

## Limitations

Review-informed heuristic, not definitive taxonomy; single-author audit trail; Anglophone/conceptual-weighted corpus; AI-aware validity captures an emerging debate, not settled evidence; should not yet be used as a formal evaluation instrument. Future research: module-redesign tests, inter-rater agreement on the operational tool, student experience of the six dimensions, comparative cross-disciplinary studies.
