Concept
Evaluative Judgement
Evaluative judgement — the capacity to make sound judgements about the quality of one's own work and the work of others, against criteria one can reason about rather than recite. In the generative AI era it has moved from a desirable graduate attribute to a load-bearing capability: when a tool can produce plausible finished work, the capability that distinguishes a competent learner is the ability to appraise that output — to judge what is good, what is wrong, what to accept, what to reject, and why. This makes evaluative judgement a legitimate object of Assessment in its own right, and the practical pivot for moving institutions from detection toward validity-centred and authentic assessment design.
Questions to Consider
- Think of something you can judge well. How did you learn that judgement — explicitly taught, or absorbed by making and comparing — and what does that imply for how it should be assessed?
- A rubric can be applied mechanically or understood. If a student follows criteria correctly but cannot say why one piece of work beats another, which capability has the assessment actually measured?
- When a language model drafts a plausible solution, the learner's remaining work is largely judgement: what to keep, what to fix, what to discard. Is that a diminished task, or a more demanding one than producing the draft?
- A student can accept an AI suggestion because it reads well, or because it survives their own scrutiny. From the outside the submission may look identical — what would let an assessor see the difference?
- A multisite experiment found that direct AI feedback produced the largest revision gains but risked passive outsourcing of judgement. What does that trade-off suggest about how you would design feedback to build rather than bypass judgement?
- Peer and AI review can train calibration, but both can become performative. How would you design assessment so that exercising judgement is required rather than performed?
Introduction
Assessments that ask students to produce work presume they can tell good work from poor work, first in others and eventually in their own. That presumption is no longer reliable: when generative AI can generate competent prose, code, and analysis on demand, the ability to produce a product is a weak signal of learning, while the ability to judge a product is a strong one. Evaluative judgement names that ability, and the knowledge base's assessment scholarship keeps arriving at it when asking what remains assessable and worth assessing. It sits where Feedback, Formative Assessment, and Self Regulated Learning meet, because judgement is developed by comparing work against standards and acting on the difference, and it is what makes genuinely critical engagement with AI output possible rather than aspirational.
What the construct is
Evaluative judgement is the capacity to appraise quality — one's own work, peers' work, and increasingly the work of an AI system — using criteria one can reason about rather than recite. Three features distinguish it from adjacent ideas.
- It is about quality, not correctness. Answering correctly requires knowledge; judging whether an answer is good requires standards a generator cannot supply on the learner's behalf.
- It is developed, not transmitted. Judgement grows through repeated acts of comparison — against exemplars, explicit criteria, and peers' differing approaches — which is why exemplars, calibration exercises, and peer review are its natural pedagogies.
- It is domain-situated. It is exercised inside a discipline's standards of evidence and argument, so it cannot be assessed generically any more than transfer can be assumed.
It is closely related to, but narrower than, authenticity in assessment: authentic assessment asks whether a task resembles worthwhile real-world work; evaluative judgement asks whether the learner can tell good work from poor work.
Why generative AI made it central
Four arguments converge.
- The product stops being evidence. Completion of a task no longer demonstrates the capability the task was designed to certify, because a tool can complete much of it. What remains is the reasoning that produces, checks, and accounts for the product — judgement, in short.
- It is the capability that governs AI use. Deciding whether an AI suggestion is worth accepting, adapting, or rejecting is evaluative judgement applied to machine output. Courses that cannot make it visible cannot tell responsible use from substitution, which is why integrity scholars increasingly treat critical AI use as an assessable outcome rather than a rule to be policed.
- It is the pivot from enforcement to design. Where prohibited use cannot be detected, Teichmann (2026) names cultivating evaluative judgement among the three design strands that relocate institutions from detection-led enforcement to validity-centred assessment — because where the institution cannot detect, the capability that survives inspection is whether the student can account for the work and its quality.
- It converts disclosure from confession to evidence. When a task asks students to explain what they used AI for, which suggestions they accepted or rejected, and how the final submission reflects their own judgement, the declaration becomes a demonstration of judgement rather than an admission to be interpreted punitively — which is what lowers the cost of honesty that otherwise drives disclosure underground.
A further layer concerns the source of a judgement, not only its content. AlGhamdi (2026) gives that layer a sharper vocabulary: his 13 Saudi computing students rendered two independent judgements within a single evaluation event — whether the feedback was useful, and whether its source held authority to grade. They all accepted ChatGPT's rubric-based feedback as clear and specific while overwhelmingly reserving the grade for the instructor ("I agree with it, but with one condition, that the doctor checks the feedback"), suggesting evaluative judgement in AI-mediated assessment must now include reasoning about evaluative authority itself, not only about the quality of the feedback received.
Evidence: how judgement develops, and what AI does to it
The knowledge base's collaborative and assessment literature supplies convergent evidence that evaluative judgement is learnable, displaceable, and designable — and that AI cuts both ways.
- Comparison-based training measurably builds it. In a 14-week study with 28 pre-service teachers, AI-generated strong, average, and weak exemplars anchored iterative comparison of students' drafts against reference points; evaluation focus expanded from surface language features to content, organisation, and coherence — but the reasoning stayed thin, with example-based justification doubling while the most sophisticated comparative reasoning remained rare. Judgement can be scaffolded into existence; it does not appear on demand, and the quality of the reference points matters.
- Feedback design decides whether agency survives. A multisite, cluster-randomized field experiment with 1,176 first-year undergraduates across 48 sections compared four feedback conditions for scientific argumentation: peer-only, direct GenAI, reflective GenAI (self-evaluate then critique), and hybrid (self-evaluate + peer + GenAI). The hybrid condition produced the largest argument-quality gain; direct GenAI feedback risked passive uptake — students outsourcing evaluative judgement to the system — while reflective and hybrid designs preserved epistemic agency by forcing the student to evaluate their own work first. The authors' core finding is that GenAI's educational value depends less on AI access than on whether the feedback environment preserves agency, judgement, and ownership during revision.
- Structured peer-and-AI review scales calibration. The PAIRR model (Sperber et al., 2025) combines peer review with AI review and structured reflection, and was tested in the largest study of students' use of AI feedback to date (654 students, ten writing courses). It treats the comparison between one's own judgement, peers', and AI's as the training ground, positioning evaluative judgement as an explicit learning outcome rather than an implicit by-product.
- AI feedback helps only when literacy already exists. A conceptual framework from feedback-literacy leaders (Boud, Dawson & Yan) analyses feedback across eliciting, processing, and enacting, using two contrasting IELTS-with-ChatGPT cases: a student with low feedback literacy used a vague prompt, received generic output, and trusted or over-copied it, while a literate student used AI critically and learned. GenAI lowers the cognitive and emotional barriers to seeking feedback but can itself be hallucinated, biased, or generic — which is why teachers are argued to model evaluative planning and train judgement, not just prompting.
- Sustainable judgement, not momentary feedback, is the gap. A scoping review of authentic assessment (Zhan, Boud & Du) found AI-formative feedback abundant but sustainable feedback — transferable to future contexts — present in only 4 of 23 formative studies. Feedback that closes the current task while leaving the learner unable to judge future work is the evaluative-judgement failure in its clearest form.
- Assessment can be designed to reward judgement instead of output. The response-region model formalises why: redesign lowers the payoff from hidden outsourcing and raises the value of visible reasoning, explanation, and critique. Asking students to explain which AI suggestions they accepted and rejected turns the decision itself into assessable evidence.
How misuse displaces judgement
The clearest risk is displacement, and it is evidenced in the feedback literature rather than only theorised. When the tool supplies the rubric, the feedback, and the monitoring, the learner's metacognitive and evaluative practice is displaced rather than supported — the assessment becomes a performance of judgement the student never exercised, the same failure mode as cognitive offloading expressed in the assessment register. Direct GenAI feedback in the multisite experiment pushed toward exactly this passive outsourcing. Grading pressure compounds it: where self-assessment or reflection on AI use is itself graded, students can produce the reflection the rubric wants rather than the reasoning, so genuine judgement requires making the exercise consequential — a decision that shows up in later work, an oral explanation, or a documented rationale that is itself examined.
Implications for practice
- State it as an outcome. If appraising work quality — including AI output — is what matters, write it into the learning outcomes and assess it, rather than leaving it as a by-product of the task.
- Assess the decision, not only the artefact. Ask students to justify what they kept, changed, or rejected; this turns the thinking into evidence and makes substitution visible without surveillance.
- Use exemplars and calibration deliberately. Comparative judgement against strong, average, and weak exemplars is the mechanism the evidence supports, and AI makes generating those exemplars cheap — one of its clearest pedagogical uses.
- Prefer hybrid and reflective feedback over direct output. Self-evaluation before AI critique, or self + peer + AI, preserves the Agency that direct AI feedback erodes; design the feedback environment rather than just adding a tool.
- Build toward sustainable judgement, not momentary fixes. Feedback that cannot transfer to the next task trains nothing durable.
- Pair it with process and orals. Process artefacts and short oral explanations show judgement in action where a final product cannot, which is a core reason AI-era assessment redesign relies on them.
Connected Concepts
- AI Education — AI in education (umbrella)
- Assessment Validity
- Feedback
- Feedback Literacy
- Formative Assessment
- Authentic Assessment
- Academic Integrity
- Self Regulated Learning
- Metacognition
- Critical Thinking
- Assessment
- AI Detection
- Agency
- Higher Ed
Connected Articles
- AI Internal Feedback Evaluative Judgments — How AI-supported internal feedback develops evaluative judgments, and where reasoning stays thin
- GenAI Feedback Design Multisite Experiment — Hybrid self/peer/AI feedback preserves agency and judgement better than direct AI feedback
- Pairr AI Peer Review 2025 — Peer-and-AI review with structured reflection as calibration at scale
- Zhan Boud Dawson GenAI Feedback Engagement — How feedback literacy governs whether AI feedback teaches or substitutes
- Zhan Boud Du Authentic Assessment Scoping Review 2025 — The sustainability gap in AI feedback
- Teichmann Detecting Undetectable Misconduct 2026 — Evaluative judgement as one of the design strands replacing detection-led enforcement
- Mohamed Temimi Assessment Imperfect Information Disclosure 2026 — Making visible reasoning and judgement the thing the rubric rewards
- Learner Centered Feedback AI — Teacher evaluative judgement with AI feedback tools
- Beyond Detection Authentic Assessment AI 2025 — Beyond detection: authenticity redesigned rather than policed
- Coauthorship Integrity Reconceptualising Assessment Validity For The Age Of Gene — Reconceptualising assessment validity for the age of generative AI
- Du Yuan Epistemic Dependence 2026 — Contestability, recoverability, and the criteria separating reliance from dependence
- Rethinking AI Writing Feedback Literacy — Feedback literacy for AI-assisted writing
- Student Perspectives AI Writing Grading 2026 — Who Should Grade My Work? Student Perspectives on Transparent AI-Assisted Writing Assessment in Higher Education