On this page

Automated Essay Scoring (AES) — the use of AI to evaluate and score written essays, spanning traditional statistical approaches, fine-tuned language models, and increasingly accessible Large Language Models (LLMs)-based prompting strategies. AES research in this knowledge base covers scoring accuracy, fairness and bias, psychometric validity, and practical Accessibility for educators.

Questions to Consider

  • Automated Essay Scoring has moved from traditional statistical models to LLM-based prompting. A key tension on this page is between accuracy and accessibility — fine-tuned models score well but are impractical for most educators. Which would you prioritize in your own context, and why?
  • A major finding is that simply including exemplar essays in the prompt brings LLM-human agreement close to human-human reliability — and a cheaper model can match a more expensive one. What does this suggest about how much of AES quality is the model versus how it's prompted?
  • The page warns that AI scoring can systematically underestimate students from linguistically diverse backgrounds. Before reading, if you saw an AI give a lower score to a non-native speaker's essay, would you have assumed it was a 'bias problem' or just 'the score'? What would change how you respond?
  • A self-referential approach assesses L2 writers by comparing their writing to their own prior work rather than to native-speaker norms. How does the choice of comparison baseline change what a score means — and which students might it treat more fairly?
  • AES intersects with formative assessment when used for feedback rather than grading. When would a machine's feedback on an essay be genuinely useful to a developing writer, and when might it flatten the kinds of qualitative feedback a human editor would give?

Introduction

Automated Essay Scoring has a long history in educational technology, from early statistical models to modern LLM-based approaches that can evaluate essays holistically without large pre-scored datasets. The key tension in AES research is between accuracy and accessibility — while fine-tuned models achieve strong results, they are resource-intensive and impractical for most educators.

  • Zhang et al. RACES uses reward alignment to make LLM essay scoring both accurate and consistent, addressing a core AES validity concern.

Key research themes

Prompting-based AES has emerged as the most accessible approach. The Choi et al. anchor paper study shows that including exemplar essays in prompts brings LLM-human agreement close to human-human reliability, with GPT-4o mini achieving comparable results to GPT-4o at lower cost. This connects to broader Prompt Engineering research and makes AES feasible for teacher use.

Psychometric and trait-level scoring moves beyond holistic scores. PsyScore provides a psychometrically-aware framework for trait-adaptive scoring with ZPD-grounded feedback. ICLE++ models fine-grained traits for holistic essay scoring, advancing the precision of automated evaluation.

Bias and fairness is a critical concern. Feser & Tschisgale found that AI scoring systematically underestimates students from linguistically diverse backgrounds, highlighting the need for Bias Mitigation and Equity considerations in AES deployment.

L2 and self-referential assessment explores non-native writing contexts. Profile-based L2 assessment uses a self-referential approach comparing student writing to their own prior work rather than native-speaker norms.

Interpretability and feature weighting opens the blackbox of how LLMs actually score. Wang et al. (2026) compared three LLMs (Qwen, GPT, Gemini) with human raters on non-native English essays across sixteen textual features, finding strong overall alignment but distinct weighting: LLMs emphasized grammatical accuracy, lexical sophistication, and syntactic complexity, while human raters prioritized content completeness and visual presentation. Critically, LLMs shifted their weighting by proficiency level — placing more weight on language errors for low-proficiency students and increasingly rewarding linguistic sophistication for high-proficiency students — whereas human raters maintained a more stable framework. Liu, Ye, and Yan (2026) extend this with a five-model evaluation framework (GPT-4.1, Llama 4 Maverick, Gemini 2.5 Flash, Claude Sonnet 4, DeepSeek R1) on 60 long essays, using causal discovery to reveal distinct evaluative heuristics: most models prioritized lexical precision and fluency, while others emphasized syntactic complexity or cross-domain integration, and some showed inconsistency, score compression, or systematic underestimation. Together these studies establish that AES validity depends not only on overall agreement but on how models weight features and whether that weighting is stable across learner subgroups — directly informing Assessment Validity and Bias Mitigation auditing.

Item-type boundaries and the limits of essay grading. In a mixed-format university exam, Falahat, Das, Bhaumik & Thambi (2026) found ChatGPT-5's concordance with faculty was substantial-to-near-perfect on objective items (CCC 0.935–1.000) but dropped sharply on open-ended responses — near-zero to negative for short-answer and only 0.341–0.854 for essay questions — and a structured rubric did not consistently improve essay agreement. This bounds AES validity: model fluency helps on well-specified items but does not carry over to holistic essay scoring, where contextual interpretation of partial-credit responses still favors human judgment.

Issue-type boundaries in diagnostic agreement. Yao and Fan (2026) used DeepSeek-R1 to label specific writing issues before revision and compared its labels with teacher diagnosis across three cycles (mean F1 = 0.768, SD = 0.037). Agreement was highest on locally cued issues such as vocabulary word choice (F1 = 0.824) and lowest where the issue turns on a claim–evidence relation (source use and evidence, 0.579), so agreement in a diagnostic role depends on the issue type, not only on the model or the prompt.

High-stakes deployment evidence. Field evidence from a real high-stakes deployment (Uruguay's Acredita EB, 2024–2025; Curi et al.) shows prompt-engineered GPT-5 reaching 60–80% agreement with trained human raters across a 15-item analytic Spanish rubric, roughly five percentage points below human inter-rater agreement for most items and never more than 15 points below, with run-to-run consistency above 90% for almost all items. Vocabulary, syntax and spelling were the weakest dimensions, and the spelling item had to be handed to a deterministic grammar checker (LanguageTool, 68% accuracy, 100% consistency) because token-level orthography is where the model is least stable. Prompts built for one exam edition transferred to the next with only topic-specific edits, and the AI was systematically stricter than humans — under-grading rather than over-grading. For AES design, the lesson is that prompting-based scoring can approach human agreement even in a national exam, while its remaining weakness sits at the level of low-level language conventions rather than holistic writing quality.

The human Benchmark, and what it can and cannot license. The OpRaise study is the strongest test in this knowledge base of whether LLM marking is ready for routine use, and its answer turns on the benchmark rather than the model. It compared three frontier systems (Claude Opus 4.6, GPT-5.4, Gemini 3 Flash), each under 27 prompt configurations crossing rubric specificity, calibration and scoring strategy, against the moderated marks of 761 authentic Psychology essays from 125 students at three UK universities. Human marks were adopted as ground truth because academic judgment is the socially accepted standard, and the authors note that human markers agree only moderately with each other — which bounds how strong AI–human agreement could reasonably be demanded to be. Against that benchmark, agreement on the UK degree band ranged from 35 to 65 percent by institution (63 percent at Cambridge, 53 percent at Nottingham, 35 percent at Manchester Metropolitan) and did not transfer between them, so the report's central recommendation is local validation on an institution's own Assessment materials. Two findings generalize beyond this corpus. First, reliability and agreement pull apart: every model re-marked nearly identically (ICC 1.00, 1.00, 0.97) and the models agreed with each other more closely than with humans (three-model ICC 0.91), yet all three agreed on the degree band for only 56 percent of submissions — self-consistency is not validity. Second, AI marks compressed toward the middle of the scale (a compression score of 0.47–0.82, with the crossover where AI and human agree on average sitting in the upper 50s to low 60s), making AI least accurate precisely at the grade boundaries that separate Firsts from Upper Seconds and passes from fails. Vocabulary range, connectives, sentence complexity and text length predicted AI marks with small but significant effects while their relationship with human marks was broadly negligible — a direct demonstration of the linguistic-bias concern, and one that no prompting strategy tested removed.

AES as a training signal rather than a judge. Scoring models are usually studied as evaluators, but SWIM uses one as a reward function: Do, Kontak and Sachan (2026) freeze a multi-trait AES verifier and score it against generated student essays, turning the predicted trait profile into a dense per-essay reward (the mean trait-normalized distance from the target profile) for GRPO, deliberately not the Quadratic Weighted Kappa metric itself, because QWK is defined over a batch of target-prediction pairs and is not a per-sample signal, while exact-match rewards are too sparse in the multi-trait setting. The gains were re-checked against a DeBERTa scorer and an evaluation-only verifier the policy never trained against (0.598 versus 0.479 for SFT, and 0.647 versus 0.501), which is the check that separates genuine proficiency control from adaptation to the reward model. For AES research this is a second validity demand: a scorer used as a training target is being optimized against, so its own trait weighting and language biases propagate into everything the generator learns - the same feature-weighting asymmetries (grammatical accuracy, lexical sophistication and syntactic complexity weighted more heavily by models than by human raters) documented elsewhere on this page, now shaping training rather than only marks.

AES sits at the intersection of Automated Assessment, Writing, and Generative AI. It connects to Formative Assessment when used for feedback rather than grading, to Feedback Loop when integrated into iterative writing processes, and to AI Literacy when educators understand and calibrate AES tools. The Assessment Validity and Educational Measurement concepts are essential for ensuring AES scores are meaningful and fair.

  • Agreement, error and what a hybrid scorer adds. In a small open-ended marketing-writing corpus the LLM out-scored both deterministic rules and an equal-weight hybrid on absolute agreement with human raters (ICC(2,1) .435 versus .266 and .091), with the hybrid significantly worse than the LLM alone, while score dispersion and a single near-empty response showed how strongly such estimates depend on corpus composition (Agreement and error in automated scoring of student marketing posts).

Connected Concepts

Connected Articles

Embed this page

Copy the code below to embed a chromeless version of this page in a learning management system or other website. The embedded view hides the site header, navigation, and footer.