CoTAL: Human-in-the-Loop Prompt Engineering for Generalizable Formative Assessment Scoring and Feedback

Created: 2026-08-03 | Tags: formative-assessmentautomated-gradinghuman-in-the-loopprompt-engineeringbenchmarkai-ed-evaluation
๐Ÿ“„ Full text: arXiv:2504.02323 ยท local
Cohn, Ashwin T S, Mohammed & Biswas (2026) introduce CoTAL (Chain-of-Thought Prompting + Active Learning): an LLM grading pipeline that couples Evidence-Centered Design with human-in-the-loop prompt engineering and iterative teacher/student feedback refinement. It improves GPT-4's scoring by up to 38.9% over a non-prompt-engineered baseline and generalises across science, computing, and engineering โ€” direct evidence that prompt-engineering quality, not model choice, is often the binding constraint in automated-grading.

How it works

1. Evidence-Centered Design (ECD) โ€” assessments and rubrics aligned to curriculum goals from the start 2. Human-in-the-loop prompt engineering โ€” labelled examples and prompts refined iteratively with educators 3. Chain-of-thought (CoT) prompting + active learning โ€” teacher and student feedback loops refine questions, rubrics, and LLM prompts across iterations

Findings

Connections to the wiki

Related Pages

Sources