๐ Full text: arXiv:2504.02323 ยท local
Cohn, Ashwin T S, Mohammed & Biswas (2026) introduce CoTAL (Chain-of-Thought Prompting + Active Learning): an LLM grading pipeline that couples Evidence-Centered Design with human-in-the-loop prompt engineering and iterative teacher/student feedback refinement. It improves GPT-4's scoring by up to 38.9% over a non-prompt-engineered baseline and generalises across science, computing, and engineering โ direct evidence that prompt-engineering quality, not model choice, is often the binding constraint in automated-grading.
How it works
1. Evidence-Centered Design (ECD) โ assessments and rubrics aligned to curriculum goals from the start 2. Human-in-the-loop prompt engineering โ labelled examples and prompts refined iteratively with educators 3. Chain-of-thought (CoT) prompting + active learning โ teacher and student feedback loops refine questions, rubrics, and LLM prompts across iterations
Findings
- Up to +38.9% scoring performance over a non-prompt-engineered baseline (no labelled examples, no CoT, no iterative refinement)
- Gains demonstrated across domains: science, computing, engineering (the generalisation question most grading papers ignore)
- Teachers and students rate CoTAL effective at scoring and explaining responses
- Their feedback yields insights that improve grading accuracy and explanation quality
Connections to the wiki
- Strong evidence for the prompt-engineering pipeline side of automated-grading and automated-essay-scoring literature
- ECD alignment answers the assessment-validity critique: rubrics derived from curriculum goals rather than model convenience
- Human-in-the-loop refinement is the human-in-the-loop pattern applied to grading infrastructure
- The iterative teacher-feedback loop connects to formative-assessment and feedback-loop design
- Contrasts with benchmark-first evaluation: this is about operational scoring quality, cf. benchmark and ground-truth-reliability-aied
Related Pages
- formative-assessment โ the assessment function being automated
- automated-grading โ the system category CoTAL advances
- human-in-the-loop โ the refinement architecture
- prompt-engineering โ the core technique
- benchmark โ cross-domain evaluation practice
- ai-ed-evaluation โ how AI assessment tools should be evaluated
- assessment-validity โ ECD alignment as a validity strategy
- feedback-loop โ iterative refinement with teacher/student input
- ground-truth-reliability-aied โ scoring reliability concerns
Sources
- Cohn, C., Ashwin T S, Mohammed, N., & Biswas, G. (2026). CoTAL: Human-in-the-Loop Prompt Engineering for Generalizable Formative Assessment Scoring and Feedback. arXiv:2504.02323. Under review, Computers and Education: Artificial Intelligence.