On this page

Synthesis: Generative AI has not created the need for authentic assessment — it has made weaknesses in assessment design harder to ignore. Polished products can now be generated or substantially mediated by tools, so product resemblance is an increasingly unreliable signal of capability. Tsiligkiris calls this risk construct substitution: an AI-generated or AI-mediated product is attributed to the st

Tsiligkiris (2026) reframes authentic assessment as an evidential and validity-oriented design problem in AI-rich higher education. His central distinction — authentic products vs. authenticated processes — argues that assessment validity under generative AI depends not on realistic outputs alone but on architectures that make human judgment, verification, and responsibility visible. A systematic conceptual review of 37 sources yields a six-dimension framework for redesigning assessment briefs at module and program level.

The core argument

Generative AI has not created the need for authentic assessment — it has made weaknesses in assessment design harder to ignore. Polished products can now be generated or substantially mediated by tools, so product resemblance is an increasingly unreliable signal of capability. Tsiligkiris calls this risk construct substitution: an AI-generated or AI-mediated product is attributed to the student, producing an inference about capability that reflects the tool's performance rather than the student's learning. The evidence is concrete: fully AI-generated submissions passed through a live university examination system largely undetected (Scarfe et al., 2024), and experienced markers do not reliably detect GenAI-authored work (Kofinas, Tsay & Pike, 2025).

Crucially, the evidential question survives any AI policy: whether AI use is prohibited, permitted, or required, the assessment must still generate evidence that warrants the inference being drawn — either about unaided capability, or about the capacity to direct, critique, and take responsibility for AI-supported work.

Six-dimension framework

  1. Contextual fidelity and consequential relevance — realistic problems, roles, constraints, audiences, artifacts; credibility to disciplinary or civic practice plus an intelligible reason to care (stakeholder, public output, policy brief), not surface imitation.
  2. Cognitive demand and evaluative judgment — analyze, evaluate, synthesize, interpret ambiguity, make defensible trade-offs. Evaluative judgment (Sadler 1989; Tai et al. 2018) matters more when machines can generate plausible first drafts.
  3. Process transparency and assessment integrity — reasoning, iteration, feedback use, and verification visible enough to support warranted inference (Boud 2000; Kane 2013). Architectures: staged submissions, annotated decision rationales, oral defense, feedback-use statements, process records.
  4. Student agency and bounded choice — topic/modality/case/medium choice within a common architecture, bounded by clear standards to preserve comparability and fairness.
  5. Inclusivity and representational fairness — realism can privilege particular communication styles, professional norms, and cultural capital; needs transparent criteria, Scaffolding, equivalent routes to demonstrate achievement, and attention to whose realities are represented.
  6. AI-aware validity and ethical practice — specify how AI relates to intended outcomes, permitted uses, and the evidence supporting defensible attribution of capability; evidential requirement (does the task still capture the construct?) plus ethical requirement (disclose, critique, verify, take responsibility).

Operational tool

The framework becomes a review instrument for assessment briefs: teams examine each dimension for evidence generated, validity risks, and redesign priorities — using a 4-point indicative alignment scale (weak → partial → substantial → strong). Not every assessment must maximize all six dimensions; across a program, tasks may emphasize different ones. It complements (rather than replaces) the AI Assessment Scale by treating AI permissions as part of a defensible assessment argument, alongside fairness, agency, cognitive demand, and process evidence.

What this means for practice

  • Instructors. Rewrite briefs so they generate evidence for the inference you intend — about unaided capability, or about the capacity to direct, critique, and take responsibility for AI-supported work — whether AI is prohibited, permitted, or required.
  • Instructors. Add process evidence to product-based tasks: staged submissions, annotated decision rationales, feedback-use statements, oral defense, and process records are what make validity defensible when polished products are cheap to produce.
  • Administrators. Review assessment at module and program level with the six dimensions as a briefing instrument and the four-point indicative alignment scale, accepting that not every task must maximize all six.
  • Instructors. Scaffold high-fidelity tasks deliberately, since realism can privilege particular communication styles, professional norms, and cultural capital; supply transparent criteria and equivalent routes to demonstrate achievement.
  • Researchers. Test redesigned authentic assessments against shared outcome measures and treat the framework as a heuristic for deliberation, not a scale to be scored and filed.

Limitations

  • The framework comes from a systematic conceptual review of 37 retained sources aimed at construct clarification rather than effect-size aggregation; the evidence base is characterized as conceptually expansive but operationally inconsistent.
  • It is a single-author review with no independent screening, coding, or inter-rater reliability check, drawing on an Anglophone, concept-heavy corpus.
  • The alignment scale is explicitly an indicative heuristic (weak, partial, substantial, strong), not a psychometric instrument, so the framework should not yet be used as a formal evaluation instrument.
  • The AI-aware validity dimension captures an emerging debate rather than settled evidence, and the authors note cross-disciplinary empirical evaluation of redesigned authentic assessments with shared outcome measures remains limited.

Citation

Tsiligkiris, V. (2026). From authentic products to authenticated processes: a systematic conceptual review of authentic assessment in AI-rich higher education. Assessment & Evaluation in Higher Education.

Embed this page

Copy the code below to embed a chromeless version of this page in a learning management system or other website. The embedded view hides the site header, navigation, and footer.