On this page

Summative assessment — assessment used to evaluate and certify what a learner has learned at the end of a unit, course, or program, in contrast to formative assessment which supports learning during instruction. Summative assessment typically takes the form of high-stakes examinations — written, oral, proctored, or closed-book — that assign grades, gate progression, and certify competence. In the AI era, summative assessment has become a central battleground over Academic Integrity and validity: generative AI can inflate performance on unproctored or take-home tasks, making the choice of summative format — and how it resists AI substitution — a pivotal design decision.

Questions to Consider

  • Summative assessment certifies what a student has learned at the end of a course, while formative assessment supports learning during it. Where have you seen the line between these two blur, and why might it matter that they serve different functions?
  • The page frames generative AI as reshaping summative assessment in two directions at once: AI scores exams, and students use AI to evade exam-based measurement. Which of these two pressures do you think is the bigger threat to validity, and why?
  • If unproctored or take-home tasks lose validity because AI can produce the answers, what does that imply for how assessments should be designed — and what might be sacrificed in the process?
  • Research cited on the page finds LLMs do not grade essays the same way humans do. If automated scoring is fast and consistent but grades differently, is that a fairness problem, an opportunity, or both?
  • What does a high-stakes result (a grade, a credential, admission) mean if the work behind it could have been produced by AI? How would you design an assessment you could actually trust?

Introduction

Summative assessment serves a fundamentally different function from formative assessment: it measures and certifies achievement rather than guiding next steps. It includes end-of-unit tests, final examinations, standardized and high-stakes tests (e.g., entrance exams), oral defenses, and cumulative performance assessments. Because summative results carry real consequences (grades, progression, credentials, university admission), they face particular pressures in the AI era — both as targets of automated scoring and as vulnerable measures that students may seek to game using generative AI.

The AI-era stakes: validity and integrity

The knowledge base's research documents how generative AI has fundamentally reshaped the summative-assessment landscape in two directions: AI is used to score exams at scale, and AI can be used by students to evade exam-based measurement of their own learning.

  • AI as scorer. Summative assessment increasingly relies on automated scoring of exams, essays, and short answers. Research on LLM essay grading finds LLM do not grade essays the same way humans do, raising validity and fairness questions for high-stakes automated scoring. LLMs assessing student self-explanations and automatic short-answer grading explore the reliability of LLM scoring in summative contexts, while psychometrically-aware frameworks seek to keep automated scoring trustworthy and adaptive. A 296-student handwritten general-chemistry exam illustrates why selective oversight is required: a multimodal LLM's total-score agreement with TA grading was high (R² = 0.91) yet item-level reliability varied sharply by format, and false positives (AI crediting genuinely wrong answers) tend to go undetected because students rarely contest them — so a uniform "grade everything" AI scorer is not defensible for high-stakes use without confidence-based deferral to humans (Assisting the grading of a handwritten general chemistry exam with artificial intelligence).
  • Outcome agreement can survive imprecise item scoring. Against moderated official marks, a multimodal LLM grading 10,364 handwritten Physics Olympiad and university pages correlated at r = 0.93–0.96 and recovered the same five-student international team, though exact question-part agreement reached 70% — a scorer is judged by the decisions it informs (Pathak et al. (2026)).
  • AI as evasion. Because generative AI can produce answers to written questions, unproctored and take-home summative tasks lose validity: proctored, unassisted measures are essential because non-proctored performance is inflated by AI, and guardrailed (hint-not-answer) tools can eliminate the exam penalty that unguarded AI causes. Chirikov's (2026) quasi-experiment on 500,000+ grades makes the mechanism concrete: after ChatGPT's release, courses with more AI-exposed tasks saw the share of A grades rise by 13 percentage points, and the effect concentrated in homework-heavy courses (an additional 16 pp in the triple-differences estimate) — direct evidence that unproctored homework, not genuine learning gains, is where AI inflates summative outcomes.
  • AI-generated exams. A large-scale field study and EFL assessment research examine whether AI can generate high-quality exams and assessment tasks — an emerging summative-design use of AI.

AI-resistant summative formats

A key theme in the knowledge base is that summative format determines AI-resistance — the more a task requires live, in-person, individually-probed performance, the harder it is for students to substitute AI for their own learning.

High-stakes and standardized summative assessment

High-stakes summative assessment — entrance exams, standardized tests, and certification — carries outsized consequences and is a focus of AI-era concern. The generative AI learning penalty study measured outcomes on high-school (Zhongkao) and college (Gaokao) entrance exams, finding entrance-exam scores fell 18–24% after prolonged AI use. The Effortless Trap and performance-vs-learning research warn that gains on AI-assisted tasks do not transfer to unassisted high-stakes measures.

Summative vs. formative in the AI era

The knowledge base's assessment literature consistently emphasizes that Assessment is most effective when it combines formative and summative functions — but the AI era sharpens the distinction. Because AI inflates performance on low-stakes, unproctored, and process-hidden tasks, summative (especially proctored/closed-book/in-person) measures become the crucial check on whether learning actually occurred. This motivates assessment redesign that keeps authentic, AI-resistant summative tasks (oral exams, code-review interviews, proctored examinations, process-based portfolios) as the anchor of integrity while using formative assessment to support learning along the way. See Authentic Assessment for the constructive design response.

Implications for AI in education

  • Summative format is a validity and integrity lever: AI-resistant summative formats (oral, proctored, closed-book, in-person) preserve the connection between assessed performance and actual learning.
  • Proctored/unassisted measures are the reliable signal: when students use AI, unassisted summative exams — not homework — reveal genuine learning.
  • Automated scoring needs psychometric scrutiny: using LLMs to grade high-stakes exams requires evaluation of reliability, fairness, and validity, not just accuracy. Grading error is also grade-dependent: across 32 public-health exam submissions, the LLMs avoided the scale's ends — no E or F appeared in fast mode — while the best model matched the human grade exactly on 50.0% and within ±1 grade on 90.6% (Brevik et al. (2026)).
  • AI can also generate exams: AI-assisted exam and task generation is an emerging summative-design application that itself needs quality evaluation.
  • Reconsider grading purpose, not just format. Mesny, Roberge-Maltais & Galy (2026) critique traditional summative, norm-referenced grading for encouraging superficial, fragmented learning, giving students little control or transparency, harming intrinsic Motivation, fueling stress and anxiety, and perpetuating inequities while largely assessing recall rather than real-world application. They position reassessment, standards-based grading, and ungrading as grading-focused innovations that can soften summative-heavy practice, while acknowledging these remain marginal in management education because of normative barriers — grading on a curve, external signaling (rankings, internships, accreditation), and students' instrumental mindset — and recommend incremental experimentation with institutional support.

Connected Concepts

Connected Articles

Embed this page

Copy the code below to embed a chromeless version of this page in a learning management system or other website. The embedded view hides the site header, navigation, and footer.