Romina Mahinpei, Victoria Dean, Ruth Fong, Lydia T. Liu, Manoel Horta Ribeiro (2026) โ arXiv. ๐ Full text (arXiv)
This field experiment shows that AI-generated feedback drafts can measurably increase the rate and length of feedback that teaching assistants actually deliver to students, without sacrificing perceived usefulness or instructor time-efficiency. The design keeps humans firmly in the loop: TAs could edit or discard every draft, and the intervention still produced significant gains. The finding connects directly to automated-grading and ai-feedback-quality debates: AI does not replace the grader here, but reduces the activation energy for starting a feedback document. TAs treated drafts as editable scaffolds rather than authority, which aligns with teacher-role research on maintaining instructor agency. The +10.8pp provision effect and +39.8-char length increase are rare quantified benchmarks for discretionary AI assistance in education. Because the study measured usefulness ratings alongside quantity, it also informs the feedback-loop literature: more feedback is not automatically better feedback, yet the null result on student ratings suggests the drafts did not degrade quality. The mixed-methods design bridges efficacy-study and RCT traditions in AIED evaluation. Practical implication: if deployed at scale, AI feedback scaffolding could be especially valuable in large-enrollment higher-ed courses where TA time is scarce but personalized feedback is pedagogically important. Future work should test whether gains persist across semesters and whether different subject domains moderate the effect.
Related Pages
- automated-grading โ AI-generated feedback as scaffolded output
- ai-feedback-quality โ Quality benchmarks for AI tutoring feedback
- teacher-role โ Teacher orchestration of AI classroom tools
- student-ai-interaction โ Scaling classroom student-AI dialogue
- higher-ed โ Higher education feedback provision context
- feedback-loop โ Discretionary feedback as high-value loop
- ai-assessment-human-tutors -- AI-driven assessment of human tutor training performance correlates with real-li...