AI Ed Wiki logoAI Ed WikiUse with AI

Synthesis: Mathew, Taher, Kundu, and Barbosa (2026) evaluate how out-of-the-box LLMs (GPT and Llama families) score essays compared with human graders on ASAP and DREsS datasets. Agreement with human scores is weak and varies systematically with essay quality: LLMs assign higher scores to short/underdeveloped essays but lower scores to longer essays with minor surface errors. LLM scores are internally consistent with LLM feedback, but the models rely on signals that differ from human raters. The authors conclude LLMs are not yet suitable as standalone summative graders, but are useful as a formative first-pass Feedback tool — with important Privacy/consent caveats.

The Core Finding: Weak Agreement with Systematic Bias

Across several GPT (3.5/4/5) and Llama (2/3/4) models, agreement between LLM-generated scores and human grades was generally weak (inter-rater QWK ~0.17–0.28, vs ~0.72 between two human raters). Critically, disagreement is not random — it varies systematically with essay characteristics:

  • Short/underdeveloped essays tend to receive higher LLM scores when prompt relevance and surface readability are present, even at the expense of depth of argumentation.
  • Longer, strong essays with minor surface errors tend to receive lower LLM scores — LLMs penalize grammatical/spelling mistakes that human raters tolerate when content and argumentation are strong.
  • Central tendency bias: LLMs cluster around the middle of the grading scale, avoiding extreme scores rather than using the full range.

Internal Consistency of Scores and Feedback

LLM-generated scores are generally consistent with the feedback they produce: essays receiving more praise tend to receive higher scores, and essays receiving more criticism tend to receive lower scores. This suggests models follow coherent internal criteria when evaluating essays — even though those criteria differ from human raters' implicit standards. Different LLMs also rely on different subsets of rubric traits, partly explaining the variability in grading behavior across models.

Implications for Assessment Practice

  1. Not suitable for high-stakes summative grading. Systematic differences mean strong writers may be unfairly penalized while underdeveloped work goes undetected — a significant risk where grades carry direct consequences.
  2. A useful formative first-pass tool. The coherence between scores and feedback supports a role in guiding student revision: generating Feedback on drafts, flagging surface errors, and identifying essays needing closer attention — while keeping the human as the final judge.
  3. Prioritize Feedback over the numeric score. Since LLM scores are less reliable at the extremes, educators should favor the generated feedback and apply simple calibration strategies (e.g., adjusting for essay length or surface errors).
  4. Privacy and consent are essential. Processing student essays via commercial APIs may violate FERPA/GDPR; student data could be used for model training or later flagged as AI-generated by plagiarism detectors. Open-source models running locally are a more Privacy-preserving option. Students should be informed when LLMs are involved in evaluating their work.

g their work.

Connected Concepts

Connected Articles

Citation

Mathew, J. G., Taher, S., Kundu, A., & Barbosa, D. (2026). LLMs Do Not Grade Essays Like Humans. Computers and Education: Artificial Intelligence.