LLM-Generated Feedback in Introductory Programming: A Classroom Study

Created: 2026-06-09 | Tags: llmformative-assessmentfeedback-loopstem-educationhigher-edautomated-grading

Heickal & Lan (2026) โ€” University of Massachusetts Amherst. ๐Ÿ“„ Full text (arXiv)

Presents a large-scale classroom study (N=215 students, 6,693 submissions across 17 labs) deploying AI-generated feedback through a randomized protocol in an introductory Python programming course. Students received one of three conditions: natural language hints, AI-generated failing test cases, or no AI feedback (control). The resulting dataset, ProgFeed, captures fine-grained temporal learning trajectories.

Key findings: Natural language feedback is significantly associated with higher completion rates and faster convergence to correct solutions. Test case feedback shows heterogeneous effects that depend critically on feedback validity. The form of AI-generated feedback matters โ€” evaluating feedback quality, not just its presence, is essential for understanding pedagogical impact.

This study provides one of the largest empirical validations of LLM-based automated feedback in authentic programming classrooms, with direct implications for automated grading systems and formative assessment design in CS education.

Related Pages

Citation

APA: Heickal, H., & Lan, A. (2026). A Classroom Study of LLM-Generated Feedback Intervention in Introductory Programming. arXiv:2606.08807. Accepted at IRAISE 2026.