On this page

Synthesis: Can ChatGPT Replace the Teacher in Assessment? A Review of Research on the Use of Large Language Models in Grading and Providing Feedback — This PRISMA-guided systematic review of 42 empirical studies (2023–2025) evaluates whether large language models (LLMs) can replace teachers in Assessment and Feedback. The authors conclude that LLMs match human raters on short, well-structured tasks with detailed rubrics, but cannot fully replace human judgment on complex, open-ended, or subjective work, recommending a human-in-the-loop hybrid model. Prompt quality, rubric detail, model version, and assessment language are the dominant determinants of grading and feedback quality.

Key Findings

  • LLMs perform well on closed-ended tasks and short-answer questions, often achieving accuracy comparable to human evaluators, but struggle with complex, open-ended, or subjective assignments requiring in-depth analysis or creativity (Automated Assessment).
  • Prompt quality and the use of detailed scoring rubrics or exemplar answers significantly improve the accuracy and consistency of LLM-generated grades (Prompt Engineering).
  • The highest effectiveness is achieved in hybrid systems that combine AI-driven automatic grading with teacher oversight and verification (Human-in-the-Loop).
  • LLMs cannot fully replace teachers: performance declines for longer, multilingual, or nuanced tasks, and persistent issues include between-run inconsistency, hallucinations, and feedback that is too generic or misaligned with the assigned grade.
  • No uniform grading bias emerged across studies — models were sometimes more lenient, sometimes stricter, and often avoided extreme scores.
  • LLMs can reduce teacher workload and deliver rapid, personalized Feedback at scale, particularly in large or higher-education cohorts, while automating routine grading.

Connected Concepts

Connected Articles

Citation

Jukiewicz, M., & Wyrwa, M. (2026). Can ChatGPT Replace the Teacher in Assessment? A Review of Research on the Use of Large Language Models in Grading and Providing Feedback. Applied Sciences, 16(2), 680.

Embed this page

Copy the code below to embed a chromeless version of this page in a learning management system or other website. The embedded view hides the site header, navigation, and footer.