π·οΈ Concept
Summative Assessment
Summative assessment β assessment used to evaluate and certify what a learner has learned at the end of a unit, course, or program, in contrast to formative assessment which supports learning during instruction. Summative assessment typically takes the form of high-stakes examinations β written, oral, proctored, or closed-book β that assign grades, gate progression, and certify competence. In the AI era, summative assessment has become a central battleground over Academic Integrity and validity: generative AI can inflate performance on unproctored or take-home tasks, making the choice of summative format β and how it resists AI substitution β a pivotal design decision.
Summative assessment serves a fundamentally different function from formative assessment: it measures and certifies achievement rather than guiding next steps. It includes end-of-unit tests, final examinations, standardized and high-stakes tests (e.g., entrance exams), oral defenses, and cumulative performance assessments. Because summative results carry real consequences (grades, progression, credentials, university admission), they face particular pressures in the AI era β both as targets of automated scoring and as vulnerable measures that students may seek to game using generative AI.
The AI-era stakes: validity and integrity
The wiki's research documents how generative AI has fundamentally reshaped the summative-assessment landscape in two directions: AI is used to score exams at scale, and AI can be used by students to evade exam-based measurement of their own learning.
- AI as scorer. Summative assessment increasingly relies on automated scoring of exams, essays, and short answers. Research on LLM essay grading finds LLM do not grade essays the same way humans do, raising validity and fairness questions for high-stakes automated scoring. LLMs assessing student self-explanations and automatic short-answer grading explore the reliability of LLM scoring in summative contexts, while psychometrically-aware frameworks seek to keep automated scoring trustworthy and adaptive.
- AI as evasion. Because generative AI can produce answers to written questions, unproctored and take-home summative tasks lose validity: proctored, unassisted measures are essential because non-proctored performance is inflated by AI, and guardrailed (hint-not-answer) tools can eliminate the exam penalty that unguarded AI causes.
- AI-generated exams. A large-scale field study and EFL assessment research examine whether AI can generate high-quality exams and assessment tasks β an emerging summative-design use of AI.
AI-resistant summative formats
A key theme in the wiki is that summative format determines AI-resistance β the more a task requires live, in-person, individually-probed performance, the harder it is for students to substitute AI for their own learning.
- Oral exams and assessments. Fenton (2025) argues the oral exam is a low-tech, inherently AI-resistant summative format: its real-time, interactive dialogue tests comprehension, critical thinking, and reasoning rather than memorization, prevents students from using AI to generate and memorize answers, and mirrors professional practice. Socratic tests and code-review interviews extend this to dynamic, conversational, and interview-based summative assessment.
- Closed-book, proctored, unassisted measures. Evidence and large-scale field data show that proctored closed-book exams β not inflated homework or take-home work β are the reliable signal of actual learning when students use AI. Responsible assessment frameworks embed these unassisted measures within a validity-driven redesign.
High-stakes and standardized summative assessment
High-stakes summative assessment β entrance exams, standardized tests, and certification β carries outsized consequences and is a focus of AI-era concern. The generative AI learning penalty study measured outcomes on high-school (Zhongkao) and college (Gaokao) entrance exams, finding entrance-exam scores fell 18β24% after prolonged AI use. The Effortless Trap and performance-vs-learning research warn that gains on AI-assisted tasks do not transfer to unassisted high-stakes measures.
Summative vs. formative in the AI era
The wiki's assessment literature consistently emphasizes that Assessment is most effective when it combines formative and summative functions β but the AI era sharpens the distinction. Because AI inflates performance on low-stakes, unproctored, and process-hidden tasks, summative (especially proctored/closed-book/in-person) measures become the crucial check on whether learning actually occurred. This motivates assessment redesign that keeps authentic, AI-resistant summative tasks (oral exams, code-review interviews, proctored examinations, process-based portfolios) as the anchor of integrity while using formative assessment to support learning along the way. See Authentic Assessment for the constructive design response.
Implications for AI in education
- Summative format is a validity and integrity lever: AI-resistant summative formats (oral, proctored, closed-book, in-person) preserve the connection between assessed performance and actual learning.
- Proctored/unassisted measures are the reliable signal: when students use AI, unassisted summative exams β not homework β reveal genuine learning.
- Automated scoring needs psychometric scrutiny: using LLMs to grade high-stakes exams requires evaluation of reliability, fairness, and validity, not just accuracy.
- AI can also generate exams: AI-assisted exam and task generation is an emerging summative-design application that itself needs quality evaluation.
Connected Concepts
- Remote Proctoring
- Assessment
- Formative Assessment
- Authentic Assessment
- Automated Assessment
- Assessment Validity
- Academic Integrity
- AI Ed Evaluation
- Higher Ed
- K 12
Connected Articles
- Academic Dishonesty Automated Proctoring AI 2026
- Automated Online Exam Proctoring Decade Review 2026
- Fenton Oral Exams AI Authentic Assessment 2025 β Reconsidering oral exams as authentic, AI-resistant summative assessment
- Stromberg Generative AI Learning Penalty Secondary 2026 β The generative AI learning penalty: proctored/closed-book exam evidence
- Generative AI Reduced Study Time Math β Faster completion, less learning: proctored measures essential
- Generative AI Guardrails Harm Learning β Generative AI without guardrails harms learning
- Assessing Quality AI Generated Exams Field 2025 β Assessing the quality of AI-generated exams
- Llms Do Not Grade Essays Like Humans 2026 β LLMs do not grade essays like humans
- LLM Automated Assessment Student Self Explanations β LLMs for automated assessment of student self-explanations
- Automatic Short Answer Grading β Automatic short-answer grading
- Psyscore Essay Scoring ZPD Feedback β Psychometrically-aware trait-adaptive essay scoring
- Socratic Tests Conversational Assessment β Socratic tests: dynamic, conversational, multimodal assessment
- Code Review GenAI Cs1 β Code review interviews in CS1
- Responsible Assessment AI Era Stanford 2026 β Responsible assessment in the AI era
- Test Driven AI Assisted Learning β Test-driven AI-assisted learning
- GenAI Oop Programming Assessments 2026 β GenAI performance on object-oriented programming assessments
- Brcic Effortless Trap Productive Struggle 2026 β The Effortless Trap: productive struggle and the illusion of learning
- AI Vs Human Assessment Efl Tpck 2026 β AI-generated versus human-developed assessment tasks in EFL
- Roe Assessment Twins 2026 β Assessment twins for strengthening assessment validity in the age of GenAI (Roe, Perkins & Giray 2026)