AI Ed Wiki logoAI Ed WikiUse with AI

Summative assessment β€” assessment used to evaluate and certify what a learner has learned at the end of a unit, course, or program, in contrast to formative assessment which supports learning during instruction. Summative assessment typically takes the form of high-stakes examinations β€” written, oral, proctored, or closed-book β€” that assign grades, gate progression, and certify competence. In the AI era, summative assessment has become a central battleground over Academic Integrity and validity: generative AI can inflate performance on unproctored or take-home tasks, making the choice of summative format β€” and how it resists AI substitution β€” a pivotal design decision.

Summative assessment serves a fundamentally different function from formative assessment: it measures and certifies achievement rather than guiding next steps. It includes end-of-unit tests, final examinations, standardized and high-stakes tests (e.g., entrance exams), oral defenses, and cumulative performance assessments. Because summative results carry real consequences (grades, progression, credentials, university admission), they face particular pressures in the AI era β€” both as targets of automated scoring and as vulnerable measures that students may seek to game using generative AI.

The AI-era stakes: validity and integrity

The wiki's research documents how generative AI has fundamentally reshaped the summative-assessment landscape in two directions: AI is used to score exams at scale, and AI can be used by students to evade exam-based measurement of their own learning.

AI-resistant summative formats

A key theme in the wiki is that summative format determines AI-resistance β€” the more a task requires live, in-person, individually-probed performance, the harder it is for students to substitute AI for their own learning.

  • Oral exams and assessments. Fenton (2025) argues the oral exam is a low-tech, inherently AI-resistant summative format: its real-time, interactive dialogue tests comprehension, critical thinking, and reasoning rather than memorization, prevents students from using AI to generate and memorize answers, and mirrors professional practice. Socratic tests and code-review interviews extend this to dynamic, conversational, and interview-based summative assessment.
  • Closed-book, proctored, unassisted measures. Evidence and large-scale field data show that proctored closed-book exams β€” not inflated homework or take-home work β€” are the reliable signal of actual learning when students use AI. Responsible assessment frameworks embed these unassisted measures within a validity-driven redesign.

High-stakes and standardized summative assessment

High-stakes summative assessment β€” entrance exams, standardized tests, and certification β€” carries outsized consequences and is a focus of AI-era concern. The generative AI learning penalty study measured outcomes on high-school (Zhongkao) and college (Gaokao) entrance exams, finding entrance-exam scores fell 18–24% after prolonged AI use. The Effortless Trap and performance-vs-learning research warn that gains on AI-assisted tasks do not transfer to unassisted high-stakes measures.

Summative vs. formative in the AI era

The wiki's assessment literature consistently emphasizes that Assessment is most effective when it combines formative and summative functions β€” but the AI era sharpens the distinction. Because AI inflates performance on low-stakes, unproctored, and process-hidden tasks, summative (especially proctored/closed-book/in-person) measures become the crucial check on whether learning actually occurred. This motivates assessment redesign that keeps authentic, AI-resistant summative tasks (oral exams, code-review interviews, proctored examinations, process-based portfolios) as the anchor of integrity while using formative assessment to support learning along the way. See Authentic Assessment for the constructive design response.

Implications for AI in education

  • Summative format is a validity and integrity lever: AI-resistant summative formats (oral, proctored, closed-book, in-person) preserve the connection between assessed performance and actual learning.
  • Proctored/unassisted measures are the reliable signal: when students use AI, unassisted summative exams β€” not homework β€” reveal genuine learning.
  • Automated scoring needs psychometric scrutiny: using LLMs to grade high-stakes exams requires evaluation of reliability, fairness, and validity, not just accuracy.
  • AI can also generate exams: AI-assisted exam and task generation is an emerging summative-design application that itself needs quality evaluation.

Connected Concepts

Connected Articles