Concept
Summative Assessment
Summative assessment — assessment used to evaluate and certify what a learner has learned at the end of a unit, course, or program, in contrast to formative assessment which supports learning during instruction. Summative assessment typically takes the form of high-stakes examinations — written, oral, proctored, or closed-book — that assign grades, gate progression, and certify competence. In the AI era, summative assessment has become a central battleground over Academic Integrity and validity: generative AI can inflate performance on unproctored or take-home tasks, making the choice of summative format — and how it resists AI substitution — a pivotal design decision.
Questions to Consider
- Summative assessment certifies what a student has learned at the end of a course, while formative assessment supports learning during it. Where have you seen the line between these two blur, and why might it matter that they serve different functions?
- The page frames generative AI as reshaping summative assessment in two directions at once: AI scores exams, and students use AI to evade exam-based measurement. Which of these two pressures do you think is the bigger threat to validity, and why?
- If unproctored or take-home tasks lose validity because AI can produce the answers, what does that imply for how assessments should be designed — and what might be sacrificed in the process?
- Research cited on the page finds LLMs do not grade essays the same way humans do. If automated scoring is fast and consistent but grades differently, is that a fairness problem, an opportunity, or both?
- What does a high-stakes result (a grade, a credential, admission) mean if the work behind it could have been produced by AI? How would you design an assessment you could actually trust?
Introduction
Summative assessment serves a fundamentally different function from formative assessment: it measures and certifies achievement rather than guiding next steps. It includes end-of-unit tests, final examinations, standardized and high-stakes tests (e.g., entrance exams), oral defenses, and cumulative performance assessments. Because summative results carry real consequences (grades, progression, credentials, university admission), they face particular pressures in the AI era — both as targets of automated scoring and as vulnerable measures that students may seek to game using generative AI.
The AI-era stakes: validity and integrity
The knowledge base's research documents how generative AI has fundamentally reshaped the summative-assessment landscape in two directions: AI is used to score exams at scale, and AI can be used by students to evade exam-based measurement of their own learning.
- AI as scorer. Summative assessment increasingly relies on automated scoring of exams, essays, and short answers. Research on LLM essay grading finds LLM do not grade essays the same way humans do, raising validity and fairness questions for high-stakes automated scoring. LLMs assessing student self-explanations and automatic short-answer grading explore the reliability of LLM scoring in summative contexts, while psychometrically-aware frameworks seek to keep automated scoring trustworthy and adaptive. A 296-student handwritten general-chemistry exam illustrates why selective oversight is required: a multimodal LLM's total-score agreement with TA grading was high (R² = 0.91) yet item-level reliability varied sharply by format, and false positives (AI crediting genuinely wrong answers) tend to go undetected because students rarely contest them — so a uniform "grade everything" AI scorer is not defensible for high-stakes use without confidence-based deferral to humans (Assisting the grading of a handwritten general chemistry exam with artificial intelligence).
- Outcome agreement can survive imprecise item scoring. Against moderated official marks, a multimodal LLM grading 10,364 handwritten Physics Olympiad and university pages correlated at r = 0.93–0.96 and recovered the same five-student international team, though exact question-part agreement reached 70% — a scorer is judged by the decisions it informs (Pathak et al. (2026)).
- AI as evasion. Because generative AI can produce answers to written questions, unproctored and take-home summative tasks lose validity: proctored, unassisted measures are essential because non-proctored performance is inflated by AI, and guardrailed (hint-not-answer) tools can eliminate the exam penalty that unguarded AI causes. Chirikov's (2026) quasi-experiment on 500,000+ grades makes the mechanism concrete: after ChatGPT's release, courses with more AI-exposed tasks saw the share of A grades rise by 13 percentage points, and the effect concentrated in homework-heavy courses (an additional 16 pp in the triple-differences estimate) — direct evidence that unproctored homework, not genuine learning gains, is where AI inflates summative outcomes.
- AI-generated exams. A large-scale field study and EFL assessment research examine whether AI can generate high-quality exams and assessment tasks — an emerging summative-design use of AI.
AI-resistant summative formats
A key theme in the knowledge base is that summative format determines AI-resistance — the more a task requires live, in-person, individually-probed performance, the harder it is for students to substitute AI for their own learning.
-
Oral exams and assessments. Fenton (2025) argues the oral exam is a low-tech, inherently AI-resistant summative format: its real-time, interactive dialogue tests comprehension, critical thinking, and reasoning rather than memorization, prevents students from using AI to generate and memorize answers, and mirrors professional practice. Socratic tests and code-review interviews extend this to dynamic, conversational, and interview-based summative assessment.
-
Closed-book, proctored, unassisted measures. Evidence and large-scale field data show that proctored closed-book exams — not inflated homework or take-home work — are the reliable signal of actual learning when students use AI. Responsible assessment frameworks embed these unassisted measures within a validity-driven redesign. Where exams stay online, remote proctoring takes over that role, and the corpus's two reviews of automated proctoring find privacy and fairness concerns alongside detection gains (Ensuring Academic Integrity through Automated Online Exam Proctoring: A Decade-Long Systematic Review, A Comprehensive Review of the Changing Landscape of Academic Dishonesty in Automated Proctoring in the Era of Artificial Intelligence).
-
Pair a vulnerable task with a confirming twin. Roe, Perkins & Giray (2026) keep a GenAI-vulnerable task for its learning value but pair it with a second task assessing the same outcomes, making the mark interdependent through a confirmatory threshold or weighting so the twin certifies the result.
High-stakes and standardized summative assessment
High-stakes summative assessment — entrance exams, standardized tests, and certification — carries outsized consequences and is a focus of AI-era concern. The generative AI learning penalty study measured outcomes on high-school (Zhongkao) and college (Gaokao) entrance exams, finding entrance-exam scores fell 18–24% after prolonged AI use. The Effortless Trap and performance-vs-learning research warn that gains on AI-assisted tasks do not transfer to unassisted high-stakes measures.
Summative vs. formative in the AI era
The knowledge base's assessment literature consistently emphasizes that Assessment is most effective when it combines formative and summative functions — but the AI era sharpens the distinction. Because AI inflates performance on low-stakes, unproctored, and process-hidden tasks, summative (especially proctored/closed-book/in-person) measures become the crucial check on whether learning actually occurred. This motivates assessment redesign that keeps authentic, AI-resistant summative tasks (oral exams, code-review interviews, proctored examinations, process-based portfolios) as the anchor of integrity while using formative assessment to support learning along the way. See Authentic Assessment for the constructive design response.
Implications for AI in education
- Summative format is a validity and integrity lever: AI-resistant summative formats (oral, proctored, closed-book, in-person) preserve the connection between assessed performance and actual learning.
- Proctored/unassisted measures are the reliable signal: when students use AI, unassisted summative exams — not homework — reveal genuine learning.
- Automated scoring needs psychometric scrutiny: using LLMs to grade high-stakes exams requires evaluation of reliability, fairness, and validity, not just accuracy. Grading error is also grade-dependent: across 32 public-health exam submissions, the LLMs avoided the scale's ends — no E or F appeared in fast mode — while the best model matched the human grade exactly on 50.0% and within ±1 grade on 90.6% (Brevik et al. (2026)).
- AI can also generate exams: AI-assisted exam and task generation is an emerging summative-design application that itself needs quality evaluation.
- Reconsider grading purpose, not just format. Mesny, Roberge-Maltais & Galy (2026) critique traditional summative, norm-referenced grading for encouraging superficial, fragmented learning, giving students little control or transparency, harming intrinsic Motivation, fueling stress and anxiety, and perpetuating inequities while largely assessing recall rather than real-world application. They position reassessment, standards-based grading, and ungrading as grading-focused innovations that can soften summative-heavy practice, while acknowledging these remain marginal in management education because of normative barriers — grading on a curve, external signaling (rankings, internships, accreditation), and students' instrumental mindset — and recommend incremental experimentation with institutional support.
Connected Concepts
- Remote Proctoring
- Assessment
- Formative Assessment
- Authentic Assessment
- Automated Assessment
- Assessment Validity
- Academic Integrity
- AI Ed Evaluation
- Higher Education
- K-12
Connected Articles
-
Large language models as grading assistants in public health education: a method-comparison study of essay-style exam assessment — LLM graders compress the grade scale and avoid the extremes in high-stakes essay assessment
-
Reconsidering the Use of Oral Exams and Assessments: An Old Way to Move Into a New Future — Reconsidering oral exams as authentic, AI-resistant summative assessment
-
The Generative AI Learning Penalty: Evidence from Chinese Secondary Education — The generative AI learning penalty: proctored/closed-book exam evidence
-
Artificial Intelligence and Grade Inflation — AI task displacement as a mechanism of grade inflation; homework-heavy courses (Chirikov 2026)
-
Faster Completion, Less Learning: Generative AI Reduced Study Time on Math Problems and the Knowledge They Build — Faster completion, less learning: proctored measures essential
-
Generative AI without guardrails can harm learning: Evidence from high school mathematics — Generative AI without guardrails harms learning
-
Assessing the Quality of AI-Generated Exams: A Large-Scale Field Study — Assessing the quality of AI-generated exams
-
LLMs Do Not Grade Essays Like Humans — LLMs do not grade essays like humans
-
Exploring the Effectiveness of Using LLMs for Automated Assessment of Student Self Explanations in Programming Education — LLMs for automated assessment of student self-explanations
-
Confidence Estimation in Automatic Short Answer Grading with LLMs — Automatic short-answer grading
-
PsyScore: A Psychometrically-Aware Framework for Trait-Adaptive Essay Scoring and ZPD-Scaffolded Feedback — Psychometrically-aware trait-adaptive essay scoring
-
The Theoretical Foundation of Socratic Tests: Dynamic, Multimodal, Conversational Examinations — Socratic tests: dynamic, conversational, multimodal assessment
-
Combating Harms of Generative AI in CS1 with Code Review Interviews and a Flipped Classroom — Code review interviews in CS1
-
Responsible Assessment in the AI Era: Key Insights from a Future-Focused Conference — Responsible assessment in the AI era
-
Test-Driven, AI-Assisted Learning: Replacing Lectures with Weekly Closed-Book Tests — Test-driven AI-assisted learning
-
Revisiting the Performance of Generative Artificial Intelligence on Introductory Object-Oriented Programming — GenAI performance on object-oriented programming assessments
-
The Effortless Trap: Productive Struggle, AI, and the Illusion of Learning — The Effortless Trap: productive struggle and the illusion of learning
-
AI-Generated versus Human-Developed Assessment Tasks in EFL Context: Insights from TPCK Model — AI-generated versus human-developed assessment tasks in EFL
-
Assessment twins: An approach for strengthening assessment validity in the age of generative AI — Assessment twins for strengthening assessment validity in the age of GenAI (Roe, Perkins & Giray 2026)
-
Large Scale AI Grading of Handwritten Physics Assessments: Score Agreement and Olympiad Team Selection Outcomes — AI grading of handwritten physics assessments (Olympiad)
-
Assisting the grading of a handwritten general chemistry exam with artificial intelligence