On this page

Automated assessment — the use of AI to evaluate student work, from formative quizzes to high-stakes exams. Automated assessment spans multiple modalities — multiple-choice, short answer, essay, code, and performance-based evaluation — and ranges from direct automated grading to confidence-aware systems that report calibrated uncertainty alongside their scores.

Questions to Consider

  • Automated assessment ranges from multiple-choice scoring to essays, code, and performance-based evaluation. What's the biggest difference you'd expect between how reliably AI can grade a multiple-choice quiz versus a free-form essay — and why?
  • A core design idea here is 'confidence awareness': AI graders that report how certain they are, flagging low-confidence cases for human review rather than issuing a single unqualified score. How would a grade that said 'I'm 70% sure about this' change how you'd use or trust an automated score?
  • Research shows automated scoring can systematically disadvantage non-native speakers — the AI scores the language rather than the understanding. Why might an automated grader be especially prone to this kind of unfairness, even when it achieves high agreement with human raters overall?
  • One study found that how validation is done can inflate reported performance: a naive cross-validation method reported near-perfect results that dropped dramatically under more rigorous trial-independent validation. What does this cautionary lesson suggest about how you should read any claim that an AI assessment system 'works'?
  • The page argues calibrated confidence enables human-in-the-loop workflows, supports trust calibration, and strengthens measurement validity. If an automated system routed its most uncertain cases to a human reviewer, what would you want to know about how those cases are chosen before you trusted the split?
  • Automated grading is described as one of the most mature AIED applications, yet grading without useful feedback has limited educational value. How might a focus purely on producing scores — rather than usable feedback — change what students actually get out of an AI-graded assessment?

Introduction

Assessment modalities

Automated grading

Automated grading is one of the most mature and widely-deployed AI in education applications — AI systems that evaluate student work, from multiple-choice scoring to essay assessment and code review. Its grading modalities include:

  • Students separate feedback utility from evaluative authority: AlGhamdi (2026) documents a case where an AI score counted toward grades and students knew it — 13 computing students whose scanned handwritten work was evaluated by ChatGPT against a rubric covering clarity, organization, sentence correctness and conciseness. Every participant found the feedback useful, yet all separated that utility from authority, and several rejected the one-way, non-dialogical character of automated grading — "ChatGPT is not like people. You talk to them and explain excuses... I like normal teacher [to] grade my work." The study positions automated evaluation as acceptable for surface revision but not as a final assessor, with conditional trust mediated by instructor oversight.

  • Short answer grading: confidence-aware ASAG evaluate free-text responses, with confidence calibration critical — systems must know when grading is reliable. A scoping review of short-answer auto-marking in science (2017–early 2024) documents the field's history: Morley et al. found BERT-family models (base, RoBERTa, DistilBERT, SciBERT) dominated — used in 20 of 21 studies, peaking in 2021 — before GPT-based approaches were adopted via Prompt Engineering rather than fine-tuning from roughly 2022. Models augmented with domain data (textbooks, rubrics, further pre-training) consistently outperformed those without, yet few could justify marks in human-comprehensible terms and bias across demographic and linguistic groups was rarely examined — leading the authors to argue auto-markers should support rather than replace human examiners. A PRISMA-guided systematic review of 42 empirical grading and feedback studies (2023–2025) generalizes this caution across the field: LLMs match human raters on closed-ended tasks and short-answer questions but cannot fully replace human judgment on complex, open-ended, or subjective work requiring in-depth analysis or Creativity, and the highest grading effectiveness is achieved in hybrid systems that pair AI-driven grading with teacher oversight and verification (Can ChatGPT Replace the Teacher in Assessment? A Review of Research on the Use of Large Language Models in Grading and Providing Feedback).

  • Essay scoring: Automated Essay Scoring systems like anchor-based AES use prompting strategies to approach human-level reliability. AIAWE extends automated evaluation to broader writing assessment.

  • Code review: Linux Bash grading and CS1 code review demonstrate automated assessment in computing education.

  • A prompt persona can collapse a grader, and fine-tuning can repair it: Habibullah et al. (2026) graded a practical computer vision exam (570 dual-graded students) under 171 configurations and a second machine learning exam (1,038 students) under 162 more. A short "strict grader" preamble drove 14 of 17 open-weight models out of the graded band (MAE ≥ 8), three stopping grading altogether; the damage traced to two credit-withholding sentences rather than tone or model scale, and its direction was exam-specific, improving seven models on the second exam. Light LoRA fine-tuning on roughly 3,900 pooled graded examples brought five small open models to parity or better with a human grader, so the vulnerability is a prompting artifact rather than a capability ceiling.

  • Formative integration: A-level science automation and CoTAL show how automated grading feeds into Formative Assessment cycles.

  • Bias and fairness: Language bias in physics scoring documents how automated grading can disadvantage non-native speakers — connecting to Bias Mitigation and Equity.

  • Benchmarking GenAI models for open-ended grading: Pecuchova, Benko & Drlik (2025) benchmarked eleven GenAI and sentence-embedding models against two expert human graders on 1,885 open-ended responses to 24 software-engineering questions. Only GPTo1 reached almost-perfect agreement (Fleiss' Kappa 0.82, low false positives and false negatives across grade categories), with Claude3 and PaLM2 close behind; context-sensitive GenAI models robustly handled short, diversely-phrased student answers, while reference-based embedding models (e.g., BERT's 345 false positives) systematically penalized correct-but-divergently-worded responses — evidence that model choice, not just grading format, shapes reliability for free-text open-ended assessment.

  • Prompt exemplars decide scoring agreement: Takahashi et al. (2026) automated hierarchical diagnostic reasoning (HDR), an applied-competence task in which learners find and explain deliberately embedded errors in a short case. Using GPT-4o with no fine-tuning, the team generated new problems and scored student descriptions: generated and human-authored problems carried the same internal consistency (Cronbach's α = 0.78) and showed no significant score-distribution difference across 100 participants, and the format's correlation with a reading-comprehension control stayed low (α = 0.36 and 0.41). Scoring accuracy depended on the prompt rather than the model — without worked examples Q2 precision fell to 0.42, while adding human-scored correct and incorrect exemplars lifted agreement to 98% or higher on every question (Q1 and Q2 at 100%) — against a measured human cost of 138 minutes of expert scoring for 117 responses plus 30 minutes of consensus.

  • Multi-agent coding of teachers' open-ended responses, where the framework rather than the model carries reliability: Copur-Gencturk et al. (2026) extend automated assessment beyond student work to the coding of teachers' own content knowledge (CK) and pedagogical content knowledge (PCK), asking whether AI can reproduce the trained human coding of 268 U.S. middle school mathematics teachers' written responses about ratios and proportional relationships. A purpose-built multi-agent LLM, GradeOpt, assigns three GPT-4o-backed roles — a Grader that applies the human coding manual, a Reflector that diagnoses each disagreement with the human code, and a Refiner that folds clarification points back into the manual over up to six rounds — and matched trained human coders closely on CK (κ = .88 and .85) while reaching substantial agreement on the far more interpretive PCK items overall (κ = .68; weighted κ = .79). Two NLP baselines and a single-prompt GPT-4o model never exceeded κ = .39 or weighted κ = .55 on PCK, so the authors attribute reliability to refining instructions against the coding manual rather than to model scale, and recommend expert review wherever adjacent coding levels blur, since GradeOpt's residual errors were typically adjacent codes rather than wild misses.

  • Mixed-format exam grading (2026): Falahat, Das, Bhaumik & Thambi (2026) evaluated ChatGPT-5 against human faculty grading of a 21-item pharmacy exam (16 students) spanning multiple-choice, select-all-that-apply, fill-in-the-blank, listing, short-answer, and essay items. Concordance was substantial-to-near-perfect for objective item types (CCC 0.935–1.000 — aided by correct answers supplied during grading) but collapsed for listing (0.621–0.708), short-answer (≈0 to negative), and essay (0.341–0.854) items; providing a structured rubric did not consistently improve full-exam agreement (71.1% without vs 68.2% with). The study sharpens a methodological point that recurs across automated assessment — moderate percent accuracy frequently coexists with low concordance — so accuracy alone overstates reliability, and it recommends hybrid grading for complex, subjective, or high-stakes items.

  • Cross-family grade-band calibration on authentic marked essays (2026): Kerwat, Donaldson & Mahomed (2026) ran eight pre-specified configurations from four model families (GPT-OSS 20B, GPT-OSS 120B, Qwen 32B and Llama 3.1 8B, each at two temperatures) over the same locked cohort of 114 BAWE/Coventry essays carrying institutional Merit and Distinction labels, with one shared prompt that permitted four bands. Exact band agreement ranged from 18.4% to 54.4%: Llama 3.1 8B at temperature 0.5 led (54.4% exact, 98.2% within one band, MAE 0.474, near-zero bias) while GPT-OSS 120B was weakest (18.4%) and systematically undergraded, by an average of 1.316 bands at temperature 0.5, with Distinction accuracy trailing Merit accuracy throughout. The locked cohort contained no Pass or Fail references, so this is Merit–Distinction discrimination rather than reliability across the full grade range, and temperature was no cure — it helped one family descriptively, left another unchanged, and significantly worsened GPT-OSS 20B (39.5% to 29.8% exact). Two lessons transfer beyond the study: grade distance and directional bias reveal what agreement statistics hide, and model size is no proxy for assessment validity.

  • Rubric engineering for open-ended scoring in medical education: Olvet et al. (2026) asked whether GPT-4 could reliably score open-ended questions on pre-clerkship medical exams. After three iterations of human-driven rubric refinement at two US schools, inter-rater reliability with faculty reached substantial-to-almost-perfect agreement for three of four questions using analytic and holistic rubrics (weighted kappa up to 0.94), while the holistic-rubric item stalled at moderate (κw = 0.54). Error-pattern analysis showed discrepancies were traceable to both raters — GPT-4 over-scored when students offered multiple answers or used rubric-absent vocabulary, whereas faculty were often "overly generous" graders — leading the authors to recommend keeping humans in the loop (e.g., faculty scoring a subset to confirm accuracy). Because roughly 82% of US medical schools grade pre-clerkship work pass/fail, exact AI score agreement is often unnecessary for operational use.

  • HITL scoring for a large-scale national writing assessment (2026): Curi et al. (2026) scale prompt-based Large Language Models (LLMs) scoring to a real high-stakes Spanish writing exam (~5,000–6,000 responses/year), achieving 60–80% item accuracy and 90%+ output consistency across a 15-item analytic rubric, with a deterministic NLP checker replacing the LLM on the spelling item (100% consistency). Its distinctive contribution is decision-oriented validation over model-centric metrics: the AI's systematic under-grading bias, converted via IRT-and-Bookmark into proficiency levels, produces pass/fail discrepancies in 15.3–16.5% of cases — all of which the human-in-the-loop workflow routes to expert review, cutting responses needing full human scoring by at least 50% while preserving decision quality. This grounds hybrid grading in operational assessment rather than proof-of-concept datasets.

  • Per-dimension agreement shows where an assessor cannot detect change: SOPHIE 2.0 (Hasan et al., 2026) validated an LLM judge against the consensus of all human rating sources — standardized patients and third-party raters — on a three-dimension clinical communication rubric, reaching Pearson 0.759 and ICC(A,1) 0.746 from transcripts alone, inside the spread of the individual human raters rather than outside it. Its weakest dimension was Be Explicit (correlations of 0.414–0.598 across candidate judges), and that was also the only dimension whose scores did not move measurably across the two encounters (Δ = 0.014, p = 0.1204) — so per-dimension agreement can indicate in advance which dimensions an automated assessor is unable to show improvement on.

Confidence-aware assessment

A central design goal within automated assessment is confidence awareness: AI assessment systems that report calibrated uncertainty alongside their scores, rather than issuing a single unqualified prediction. A confidence-aware grader not only produces a grade or classification but also signals how certain it is, so that low-confidence cases can be flagged for human review and users can calibrate their Trust in the system. This is central to responsible automated assessment and connects closely to Psychometrically Aware AI and Trust Calibration.

How confidence is modeled in the knowledge base's research:

  • Fused confidence signals for short-answer grading: Confidence-Aware ASAG fuses model-based confidence signals (verbalized, latent, and consistency-based) with dataset-derived aleatoric uncertainty via Random Forest regression.
  • Confidence in multimodal student work: Confidence-aware assessment of student-drawn scientific figures extends confidence modeling to multimodal student responses.
  • Psychometric calibration of LLMs: Psychometrically aware AI advances the standard of aligning Large Language Models (LLMs) scoring with measurement theory, with calibration as a core requirement alongside Item Response Theory alignment.
  • Difficulty and response-time calibration: Programming-exam difficulty calibration repositions LLMs as auxiliary evidence sources whose difficulty estimates correlate with student pass rates.
  • Trait-adaptive essay scoring: PsyScore shows a psychometrically-aware framework can adapt essay feedback to learner traits.
  • Evaluating visual student work: DiagramIR back-translates LLM-generated math diagrams (TikZ) into an intermediate representation with deterministic checks, beating LLM-as-a-Judge on agreement with human raters and letting small models match large ones at ~10× lower cost — a scalable route to assessing non-text, diagrammatic student output.
  • Trust-gated inference in automated teacher assessment: Li, Yang & Fang (2025) place Monte Carlo dropout calibration directly in the scoring path, so that dropout variance above a learned threshold triggers a reject-and-refer output rather than a score, alongside adversarial debiasing that holds the fairness gap to 1.8% where baselines sit in the 6.4–8.2% range and an expected calibration error of 0.032 on TeacherEval-2023. Their ablation is the calibration argument in miniature: removing the trust-gated module drops inter-rater consistency to 78.6%, and the authors attribute a 41% reduction in human review workload to the gating — uncertainty handling as architecture rather than post-hoc reporting.
  • Explainability of rubric-based scoring: Bueno et al. show that model-agnostic SHAP attributions are more faithful and transferable than LLM-generated rationales for explaining rubric-based scores (e.g., classroom feedback quality), and propose deletion-based + cross-model tests as a principled way to evaluate any scoring model's explanations.

Why calibrated confidence matters:

  • Enables human-in-the-loop delegation: low-confidence cases route to a human reviewer — supporting human-in-the-loop workflows rather than blind automation.
  • Supports trust calibration: calibrated confidence lets users match their trust to the system's actual reliability, avoiding both over-trust and under-trust.
  • Improves measurement validity: confidence-aware scoring strengthens Educational Measurement and Assessment Validity by making uncertainty explicit.
  • Fairness and defensibility: flagging low-confidence cases for review reduces the risk of confidently wrong scores, especially for atypical or underrepresented responses.

  • Rubric generation and instructor-supervised grading pipelines: Mendonça et al. (2026) show that LLM-generated assessment rubrics (HARMOGEN-R) can match human-created rubrics for technical content within a ±5-point equivalence margin, with structured generation giving greater cross-model consistency. Cruz et al. (2026) evaluate an end-to-end GPT-4o grading pipeline where AI grades fell within 0.5 points of the instructor in 83% of 362 submissions (MAE 0.31) — best framed as a scalable supplement to, not a replacement for, instructor judgment.

Quality and fairness

Automated assessment quality depends on Assessment Validity and Bias Mitigation. Language bias research shows that automated scoring can systematically disadvantage certain student populations.

Deployment scenarios and what each one requires

How much alignment an institution should demand depends on what the system is allowed to decide. In a mixed-methods study of 761 authentic undergraduate essays at three UK universities, three frontier models with 27 prompt configurations each reached only 35–65 percent agreement with human markers on the degree band, and accuracy did not transfer between institutions — so the report treats candidate uses as distinct scenarios rather than one decision. Quality assurance of human marking, where AI marks in parallel and significant differences trigger human review, is the most conservative. A marking assistant, ranking or categorizing submissions by predicted quality or uncertainty, triaging complex cases, or expanding brief human comments into fuller feedback, is a genuine middle ground. AI as primary marker, with human review of a sample or of mark distributions, was judged acceptable only if system performance improves, and no stakeholder group endorsed AI as a sole marker.

Three requirements follow for any of them. Because alignment is model- and institution-specific, validation must be local and continuous — headline accuracy from elsewhere is not evidence of readiness. Because marks were compressed toward the middle, AI was least accurate at the grade boundaries and for the strongest and weakest submissions, which is where assessment decisions carry the most consequence. And because disagreement between an AI and a human mark reflects different judgment rather than mere error, discrepancies should trigger human interpretation rather than algorithmic override, with final authority over the mark retained by people. Adoption also carries AI Governance obligations the technical metrics do not capture: a right to explanation under GDPR Article 22 where automated decisions affect students, Equality Act duties once attainment level and language use are shown to influence marks, revised appeal processes for errors that are not easily caught, and an irreversibility risk — once staffing and investment shift, an institution may struggle to keep collecting the human marking it would need to return to.

Venetsanos (2026) supplies the criteria those requirements presuppose, arguing that what makes automation defensible is the epistemic status of the task rather than what the technology can technically perform. His framework limits AI to bounded verification of factual and procedural claims — all four of unambiguous documented criteria, direct comparison against established knowledge, no assessment of alternative valid approaches, and a single correct answer or pre-specified acceptable set must hold simultaneously, with ambiguity escalating to a human by default — and makes AI involvement at that level conditional on five non-negotiable principles that must all hold at once: the knowledge base must be assessor-curated and retrieval-grounded in module materials rather than the model's parametric knowledge; human assessors review every AI output before it reaches students and hold absolute override with unshared accountability; feedback must carry clear provenance and attribution; and security must be designed against adversarial use from the start, with input sanitisation for instructions hidden in white or small text, encodings, images or document metadata, since successful circumvention of an automated system spreads through student cohorts. The same paper is explicit that its principles are untested for simultaneous feasibility and that assessor time for curation, security infrastructure and high-frequency oversight may shift staff effort rather than reduce it — a caution that sits alongside the local-validation requirement above.

Security of AI-mediated grading

A red-team evaluation of an everyday grading workflow shows the attack surface that planning for adversarial use has to cover. Humble (2026) embedded five indirect prompt injections inside a synthetic essay that Microsoft Copilot (GPT-5.2) had graded fail in six of six baseline runs, iterating each across docx, pdf and htm files. Two strategies changed the grade with no visible warning — an instruction-manipulation and role-playing paragraph at 9 of 9 iterations (100%) and the same paragraph hidden behind an image layer at 17 of 18 (94%) — while a paragraph in small white text at the end of the document failed in all nine iterations, and file-metadata injections never worked, which the author attributes to metadata access being disabled in the version tested. The trust problem outlasted the exploits: after one pdf run detected an embedded instruction and stated it would grade only against the official assignment, re-running the same file raised the grade six more times with no warning, and a chat blocked by a detected attack was silently disabled rather than reported to the user. The author argues the resulting grades carry no validity claim in either direction, that an adversary needs only one working combination against layered Guardrails, and that the teacher remains the only real check on the output while the manipulation is designed not to be visible — connecting directly to the input-sanitisation and adversarial-security obligation Venetsanos sets out above.

Connections

Automated assessment connects to Assessment Validity (quality assurance), Formative Assessment (use context), Bias Mitigation and Equity (fairness), Teaching (how automation changes instructor work), and AI Feedback Quality (grading without useful feedback has limited educational value). Confidence-aware assessment is a specific mechanism within the broader agenda of Psychometrically Aware AI and a contributor to calibrated trust.

  • Large-scale AI grading of handwritten physics (2026): A multimodal model (GPT-5.5) graded 10,364 scanned pages across a national Physics Olympiad theory exam, selection camp, and university quantum-mechanics exam, achieving total-score correlations of 0.91–0.97 with official marks and recovering the same top-five Olympiad team. Revised page-by-page, evidence-location instructions improved agreement — evidence that multimodal AI can support high-stakes summative grading with careful rubric and prompt engineering (Large Scale AI Grading of Handwritten Physics Assessments: Score Agreement and Olympiad Team Selection Outcomes).

  • Selective automation for handwritten chemistry (2026): Cvengros & Kortemeyer graded a 296-student handwritten general-chemistry final page-by-page against rubric images with a multimodal, reasoning LLM, achieving high run-to-run reliability on total scores (ICC(A,1) = 0.967) and strong total-score agreement with TA grading (R² = 0.91). Because item-level reliability varied sharply by format — textual and chemical-reaction answers graded reliably while drawing and graphing scored worse than random — raw agreement was judged inadequate for high-stakes use. They operationalize human deferral through confidence filters (partial-credit thresholds, an IRT-based risk threshold, and problem-type exclusion) that convert raw AI scores into an accept/deferral policy, and show that false positives (AI crediting genuinely wrong answers) tend to go undetected since students rarely contest them — a concrete psychometrically grounded route to selective automation (Assisting the grading of a handwritten general chemistry exam with artificial intelligence).

  • Semi-open handwritten mathematics in a university exam (2026): Liu et al. graded a German undergraduate mathematics exam with GPT-4 across two extraction routes (pre-cut answer boxes versus whole pages), two optical-character-recognition routes and two rubric formats, reporting Krippendorff's alpha against human grades of 0.22-0.44 with accuracy 0.59-0.62 on the best workflows and per-problem accuracy spanning 0.27-0.91 - not acceptable for high-stakes summative use. Itemising multi-criteria grading rules generally did not improve agreement (it helped the partial-credit problem), free grading without a rubric was far worse (accuracy 0.50-0.51), and a probabilistic confidence filter meant to flag grades for human re-checking did not significantly improve accuracy while wrongly flagging 7-41 percent of responses as reliable. The study's transferable lesson is procedural: scan and transcribe the whole exam sheet with markers rather than cropping answer boxes, because students write across box borders; rewrite grading rules written for teaching assistants, since the model attempted unrequested derivations from a rule meant for a human; and note that the human-grading rules themselves paraphrased at an alpha of 0.88 among five variants, so prompt robustness is not the bottleneck.

  • LLM comparative judgment for writing screening (2026): Mercer & Reed used seven LLMs as pairwise comparative judges to score informational writing from 1,208 students in Grades 3–6 across three screening occasions. LLM-CJ scores correlated strongly with analytic rubric scores (r = .67–.73, strongest for Gemini-3.1 Pro) with AUC .82–.86 for proficiency; findings were stable across models, with little validity gain from costlier, more capable ones, while averaging three writing samples improved accuracy — sampling breadth mattered more than model choice. Predictive-bias patterns for multilingual learners matched human rubric scoring, supporting LLM-CJ as an efficient, low-cost screening approach (Validity of Large Language Model Comparative Judgment for Universal Writing Screening).

  • Quality-dependent alignment with instructor grading (2026): Comparing ChatGPT, peers, and an instructor grading the same undergraduate group projects, Usher & Faraon found ChatGPT's alignment with the instructor improved as project quality increased — its largest overestimation was for low-quality work (≈ +14 points), shrinking to ≈ +2.5 points for high-quality projects. ChatGPT also graded higher on average than both peers and the instructor, a grade-inflation tendency that undermines its reliability as a standalone summative grader, especially for weaker submissions.

  • Small-corpus agreement statistics can mislead deployment decisions. Scoring 60 marketing posts (15 students plus 15 researcher-authored low-quality anchors) against two independent human raters gave absolute agreement ICC(2,1) of .435 for the LLM, .266 for an equal-weight rule-plus-LLM hybrid and .091 for deterministic rules, with MAE of 6.28, 10.53 and 17.22 points respectively on a 0-100 scale; the hybrid was significantly worse than the LLM alone (paired ICC difference -.169, 95% CI [-.260, -.106]). Adding anchors raised inter-rater ICC from .338 to .902 and LLM agreement from .435 to .846, and a near-empty post was awarded 75 against a human mean of 30.5, showing that a single degenerate response can dominate a small evaluation. (Agreement and error in automated scoring of student marketing posts)

Connected Concepts

Connected Articles

Connected FAQs

Embed this page

Copy the code below to embed a chromeless version of this page in a learning management system or other website. The embedded view hides the site header, navigation, and footer.