1. Evidence-Centered Design (ECD) β assessments and rubrics aligned to curriculum goals from the start 2. Human-in-the-loop prompt engineering β labelled examples and prompts refined iteratively with educators 3. Chain-of-thought (CoT) prompting + active learning β teacher and student feedback loops refine questions, rubrics, and LLM prompts across iterations
CoTAL: Formative Assessment Scoring with Human-in-the-Loop Prompt Engineering
Cohn, Ashwin T S, Mohammed & Biswas (2026) introduce CoTAL (Chain-of-Thought Prompting + Active Learning): an LLM grading pipeline that couples Evidence-Centered Design with human-in-the-loop prompt engineering and iterative teacher/student feedback refinement. It improves GPT-4's scoring by up to 38.9% over a non-prompt-engineered baseline and generalises across science, computing, and engineering β direct evidence that prompt-engineering quality, not model choice, is often the binding constraint in Automated Grading.
How it works
1. Evidence-Centered Design (ECD) β assessments and rubrics aligned to curriculum goals from the start
2. Human-in-the-loop prompt engineering β labelled examples and prompts refined iteratively with educators
3. Chain-of-thought (CoT) prompting + active learning β teacher and student feedback loops refine questions, rubrics, and LLM prompts across iterations
Findings
Up to +38.9% scoring performance over a non-prompt-engineered baseline (no labelled examples, no CoT, no iterative refinement)Gains demonstrated across domains: science, computing, engineering (the generalisation question most grading papers ignore)Teachers and students rate CoTAL effective at scoring and explaining responsesTheir feedback yields insights that improve grading accuracy and explanation qualityConnected Concepts
AI Ed EvaluationAssessment ValidityAutomated GradingFormative AssessmentPrompt EngineeringLLMConnected Articles
Ground Truth Reliability AIED β Modernizing Ground Truth: Four Shifts Toward Improving Reliability and Validity in AI in EducationAaai2026 Prompting Literacy K12 β Learning to Use AI for Learning: Teaching Responsible Use of AI Chatbot to K-12 Students Through an AI Literacy ModuleAcademiclaw Student Agent Benchmark β AcademiClaw: When Students Set Challenges for AI AgentsAdaptive Pretesting Retention β Do Gains from Generative AI-Enabled Adaptive Pretesting Persist? Evidence from a Retention StudyAgent Voice Accents K12 Group Learning β Exploring How Agent Voice Accents Shape Human-AI Collaboration in K-12 Group LearningAgentic AI Education Scoping Review β Agentic AI in Education: A Scoping Review of Research Landscape, Capabilities, and the Frontier Agent ParadigmAgentic AI Pedagogical Best Practice 2026 β Agentic AI and Pedagogical Best Practice: The Tension Between Automation and LearningAgentic Workflows Education β Agentic Workflows in EducationAgents That Teach Incidental Learning β Agents That Teach: Designing Incidental Learning Back into AI-Assisted Software DevelopmentAgreement Not Quality LLM Coding Verification β Agreement Is Not Quality: Blind Expert Verification of Human and LLM Qualitative Coding When Human Consensus Is Not G...AI Adoption Training Public Sector β The Main Barrier to AI Adoption in the Public Sector is Lack of TrainingAI Agents Peer Learning Discourse β When AI Agents Teach Each Other: Discourse Patterns Resembling Peer Learning in the Moltbook CommunityAI Assessment Human Tutors β AI-Driven Assessment of Human Tutors: Linking Training Performance to Real-Life PracticeAI Assistance Discretionary Feedback β AI Assistance for Discretionary Work: Increasing Feedback Provision in Higher EducationAI Assisted Learning Modes Eeg β An exploratory behavioral and electroencephalographic study of artificial intelligence-assisted learning modes in hig...AI Availability Student Motivation β Why Put in This Much Effort?": How AI Availability Shapes Studentsβ Motivation in Introductory ProgrammingAI Campus Wellbeing Tools β AI-Driven Tools for Enhancing Campus Well-being: Prevention and InterventionAI Changing Teaching Workflows β How AI Is Changing Teaching WorkflowsAI Enabled Serious Games β AI-Enabled Serious Games: Integrating Intelligence and Adaptivity in Training SystemsAI Engineering Education Balancing Act β Using AI in engineering education: a balancing act, driven by clear purposeAI Generated Feedback Higher Ed β Artificial intelligence and feedback in university education: effectiveness and student perceptionsAI Generated Traces Novice Programmers β AI-Generated Traces for Novice Programmers: Learning Effects and Learner Differences in a Multi-Institutional StudyAI In The Wild College β AI in the Wild: A Large Scale Analysis of Authentic Interactions of College Students with Generative AIAI Interlocutor L2 Spoken Dialogue β What Changes When the Interlocutor Is an AI? Interactional Fluency and Linguistic Uptake in L2 Spoken DialogueStanford Evidence Base AI K12 2026 β AI in K-12 Evidence BaseCitation
Cohn, C., Ashwin T S, Mohammed, N., & Biswas, G. (2026). CoTAL: Human-in-the-Loop Prompt Engineering for Generalizable Formative Assessment Scoring and Feedback. arXiv:2504.02323. Under review, Computers and Education: Artificial Intelligence.