On this page

Synthesis: Wiggins (1990) proposed AA as a counterbalance to standardized tests: direct examination of "student performance on worthy intellectual tasks."

Authentic assessment (AA) has evolved from workplace-task replication toward a multi-dimensional framework encompassing professional, digital, personal, and social authenticity. The recent challenge by generative AI—which threatens the validity of any task that can be replicated à la Wiggins (1990)—makes AA's broader forms essential. Zhan, Boud & Du (2025) propose a six-dimensional design model that centers student agency and social collaboration, directly relevant to how AI assessment tools should be designed.

The Evolution of Authenticity

1990s Origins: Worthy Intellectual Tasks

Wiggins (1990) proposed AA as a counterbalance to standardized tests: direct examination of "student performance on worthy intellectual tasks."

Late 1990s HE Uptake: Workplace Replication

Joughin (1998) argued authenticity should reflect the extent to which assessment replicates professional practice or real life. This view dominated for two decades.

2020s Critique: Beyond Replication

McArthur (2023) contends AA must enable students to "influence the future and transform society" rather than merely replicate existing tasks. Ajjawi et al. (2024) broaden authenticity to contextual, task, and personal forms that reflect student experience.

Generative AI as Existential Challenge

Generative AI makes traditional workplace-replication AA newly vulnerable: any task that a language model can credibly simulate in a take-home setting loses its validity as an assessment of original student competence. The field must pivot toward forms of authenticity (digital literacy, real-time collaboration, social contribution, individual meaning-making) that AI cannot credibly counterfeit.

Six-Dimensional Framework (Zhan et al., 2025)

This scoping review of 37 empirical AA studies (2000–2024) proposes six design dimensions, moving beyond earlier frameworks (Gulikers et al., 2004; Villarroel et al., 2018):

  1. Authenticity in assessment — multiple meanings: assessment authenticity (portfolios, projects, concept maps), professional authenticity (workplace scenarios), digital authenticity (Twitter, podcasts, YouTube, LMS), self-authenticity (student identity, Well-Being), and social authenticity (citizenship, Sustainability, Ethics). Only 3 of 37 studies addressed social authenticity — a critical gap.
  2. Cognitive challenges — knowledge construction (n=29), professional skills (n=22), and 21st-century skills (n=29, led by critical thinking n=17, communication n=13). Digital literacy: only n=5.
  3. Assessment criteria — rubric use was common (n=22) but most students were passive recipients rather than co-authors. Only 3 studies co-designed rubrics with students; only 7 involved students as assessors via self/peer assessment.
  4. Feedback — formative feedback dominated (n=23), summative was common (n=12), but sustainable feedback (transferable to future contexts) appeared in only 4 studies. This mirrors the field-wide problem that AI tools also replicate: reactive, momentary feedback rather than lifelong evaluative judgment.
  5. Student agency — choices about what/how/when/where to submit appeared in only n=8 studies. Self-reflection was more common but often assigned/graded, making it potentially performative (instrumental rather than genuine).
  6. Social collaboration — mostly individual tasks (n=18) or group tasks (n=16), with few mixing both (n=3). Peer collaboration strategies (peer assessment, peer discussion) appeared in n=16 studies; teacher–student collaboration in n=16, though only 3 designed equitable teacher–student partnership (roles were usually feedback-giver, monitor, facilitator — a power imbalance); external industry/community connections in only n=5. Social construction of assessment meaning was under-theorized but present.

AI-Specific Implications

What AI Assessment Tools Get Wrong

Current AI assessment systems — MCQ generators (CODE-GEN), essay scorers (MASS), short-answer graders — focus on efficiency and standardization, replicating the very limitations Zhan et al. identify:

  • Rubric-centric: AI systems typically generate pre-defined rubrics without student co-design, replicating the "passive recipient" problem
  • Momentary feedback: AI formative feedback is abundant but rarely designed as sustainable evaluative judgment
  • Professional authenticity bias: Most AI-generated assessments simulate workplace or academic tasks, neglecting personal and social authenticity
  • Choicelessness: AI assessment systems rarely allow students to define assessment parameters, output formats, or evaluation criteria

The Four-Step Collaborative Design Framework

Zhan et al. propose a cyclical design model that AI tools could operationalize:

Step Action AI Enabler AI Risk
1. Decide goals Students + educators co-negotiate purpose and authenticity Large Language Models (LLMs)-facilitated dialogue tools Over-optimizing for what's easy to grade
2. Create context Design real-world scenarios RAG (Retrieval-Augmented Generation)-augmented scenario generation Hallucinating false domain contexts
3. Design criteria Co-design rubrics with students Collaborative rubric editors Imposing opaque algorithmic criteria
4. Plan feedback Future-oriented, sustainable feedback LLM personalization based on learner profiles Surveillance-level behavior tracking

Connections to AI Education Research

Self-Regulated Learning

Student agency in AA (choice, self-reflection, co-design) is isomorphic to the forethought → performance → self-reflection cycle. However, when self-reflection is graded, it becomes performative — students write to impress assessors rather than to learn. AI journaling tools face the same instrumentalization risk.

Metacognitive Calibration

Metacognition is required for students to evaluate their own work against co-designed rubrics. When AI provides the rubric, generates the feedback, and monitors progress, the student's metacognitive practice is displaced — the very suppression risk identified in SafeTutors and LLM Fallacy research.

Pedagogical Training

Theory-grounded training (see ISD-Agent-Bench, EduQwen) should explicitly align with the six-dimensional framework. A model trained to reward "guiding over answering" still falls short if it does not understand sustainable feedback, co-designed rubrics, or social authenticity.

Adaptive Systems

Adaptive systems that personalize only content difficulty miss the personalization of assessment authenticity. DeepTutor's multi-resolution memory and MAIC's archetype agents begin to address this, but neither incorporates student co-design of assessment parameters.

Open Questions

  1. AI-proof assessment types: Which forms of authentic assessment are robust to generative AI? In-vivo demonstrations, social contribution portfolios, co-created artifacts with auditable provenance chains, and assessments requiring real-time embodied interaction may be more resilient than take-home essays or MCQs.
  2. Student co-design at scale: Zhan et al. show co-design is rare (3/37 studies). Can AI tools enable rubric co-design at classroom or MOOC scale, or does the paradox of machine-mediated human agency undermine the authenticity itself?
  3. Sustainable feedback via LLM: Can a language model deliver feedback that students apply months later? The CDPK and ISD benchmarks test pedagogical knowledge transfer to models, not feedback sustainability transfer to students.
  4. Social authenticity deficit: Only 3 studies addressed social issues (citizenship, sustainability, ethics). How can AI assessment tools help students contribute to societal transformation rather than merely simulate it?

What this means for practice

  • Instructors. Co-design the rubric with students rather than handing one down: twenty-two studies stated the use of assessment rubrics, yet only Chang (2001), Kaya (2008), and Kearney and Perkins (2014) developed rubrics collaboratively with students, and in most cases students were passive users of the rubric.
  • Assessment designers. Build feedback students will reuse. Twenty-three studies adopted formative feedback for immediate improvement, but only four studies used sustainable feedback that empowers students to become self-directed lifelong learners — the reactive pattern that current AI feedback tools reproduce.
  • Curriculum designers. Give students real choice over what, how, when and where they take authentic assessment: only eight studies did. Pair that with ungraded self-reflection, since reflection was usually assigned as a graded task and risks performing for the assessor rather than supporting genuine self-regulation.
  • Assessment designers. Broaden authenticity beyond professional scenarios: most studies (n = 22) placed tasks in professional scenarios, six presented an authentic digital world, nine addressed students' personal experience, and only three addressed social issues such as citizenship, sustainability and ethics.
  • Curriculum designers. Design collaboration deliberately. Sixteen studies used group tasks and eighteen used individual ones; only three designed equitable collaboration between teachers and students (teacher roles were typically feedback giver, monitor or facilitator), and only five connected students with industry partners or community stakeholders.

Limitations

  • This is a scoping review of 37 empirical studies from 2000 to 2024: it maps how authentic assessment was designed, and its coding took the reported designs at face value without appraising research methodologies or study quality, so it cannot show whether any design produces the intended outcomes.
  • Western contexts are overrepresented — 27 of the included studies came from Western settings against 10 from Eastern ones, with Australia alone contributing 11, and the search covered English-language articles only, which the authors name as a limit on applicability elsewhere.
  • The underlying evidence is thin: sample sizes in the reviewed studies ranged from 5 to 493 participants, and only 29% included 100 or more.
  • Several headline gaps rest on very few studies: social authenticity on three, student choice on eight, and sustainable feedback on four.

Citation

Zhan, Y., Boud, D., & Du, Z. (2025). Designing for authentic assessment: a scoping review. Higher Education.

Embed this page

Copy the code below to embed a chromeless version of this page in a learning management system or other website. The embedded view hides the site header, navigation, and footer.