Research Article
Can ChatGPT Replace the Teacher in Assessment? A Review of Research on the Use of Large Language Models in Grading and Providing Feedback
Synthesis: Can ChatGPT Replace the Teacher in Assessment? A Review of Research on the Use of Large Language Models in Grading and Providing Feedback — This PRISMA-guided systematic review of 42 empirical studies (2023–2025) evaluates whether large language models (LLMs) can replace teachers in Assessment and Feedback. The authors conclude that LLMs match human raters on short, well-structured tasks with detailed rubrics, but cannot fully replace human judgment on complex, open-ended, or subjective work, recommending a human-in-the-loop hybrid model. Prompt quality, rubric detail, model version, and assessment language are the dominant determinants of grading and feedback quality.
Key Findings
- LLMs perform well on closed-ended tasks and short-answer questions, often achieving accuracy comparable to human evaluators, but struggle with complex, open-ended, or subjective assignments requiring in-depth analysis or creativity (Automated Assessment).
- Prompt quality and the use of detailed scoring rubrics or exemplar answers significantly improve the accuracy and consistency of LLM-generated grades (Prompt Engineering).
- The highest effectiveness is achieved in hybrid systems that combine AI-driven automatic grading with teacher oversight and verification (Human-in-the-Loop).
- LLMs cannot fully replace teachers: performance declines for longer, multilingual, or nuanced tasks, and persistent issues include between-run inconsistency, hallucinations, and feedback that is too generic or misaligned with the assigned grade.
- No uniform grading bias emerged across studies — models were sometimes more lenient, sometimes stricter, and often avoided extreme scores.
- LLMs can reduce teacher workload and deliver rapid, personalized Feedback at scale, particularly in large or higher-education cohorts, while automating routine grading.
Connected Concepts
- Feedback
- AI Feedback Quality
- Automated Assessment
- Large Language Models (LLMs)
- Teaching
- Assessment
- Formative Assessment
- Human-in-the-Loop
Connected Articles
- Can AI Evaluate Assessment? A Study of Large Language Model Meta-Assessment Performance
- Comparing Generative AI and teacher feedback: student perceptions of usefulness and trustworthiness
- Comparing GPT and human raters in essay assessment: Variability, bias, and the potential of LLM-based scoring
- LLMs Do Not Grade Essays Like Humans
- LLM-generated formative feedback in education: A qualitative systematic literature review
- Automated Grading of Open-Ended Questions in Higher Education Using GenAI Models
Citation
Jukiewicz, M., & Wyrwa, M. (2026). Can ChatGPT Replace the Teacher in Assessment? A Review of Research on the Use of Large Language Models in Grading and Providing Feedback. Applied Sciences, 16(2), 680.