Research Article
GPT-4 feedback increases student activation and learning outcomes in higher education
Synthesis: Geschwind, Graf Lambsdorff, Voss, and Hackl (2026) conduct a lab-in-the-field experiment across one semester in undergraduate macroeconomics tutorial classes at the University of Passau, comparing three feedback conditions: group-level lecturer feedback only (LF), lecturer feedback plus individual Feedback from an anonymous peer (PF), and lecturer feedback plus individual feedback from GPT-4 (AIF). Students answered eight weekly open-ended questions and received feedback across all three Hattie & Timperley dimensions (Feed-Back, Feed-Up, Feed-Forward). The authors find that AI-generated individual feedback significantly boosts student activation — sustaining the highest participation rates and producing the longest written answers across tasks — and yields the largest improvements in content learning, which they attribute to the higher reliability and quality of AI feedback relative to peers.
Key Findings
-
AIF sustains participation. Relative to lecturer-feedback baseline, students receiving GPT-4 feedback maintained the highest participation over the eight tasks (still ~50% by Task 8, vs ~30% for LF and PF); the LF-vs-AIF difference was significant at the 10% level (Fisher's exact, p = 0.054), whereas PF did not significantly outperform LF.
-
AIF induces greater effort on the intensive margin. Random-effects models on task-by-task changes show returning AIF students wrote ~29 characters more per answer than their LF counterparts (significant), while the PF effect was smaller and insignificant — evidence of stronger activation rather than mere group-level feedback.
-
AIF produces the strongest content learning gains. AIF students showed a significant improvement in content scores (~0.11, p < 0.10; rising to 0.16 when restricted to those who actually received prior feedback), while no significant differences emerged for answer style across conditions — the activation-boosting intervention also produced higher-quality output.
-
Feedback reliability drives the AI advantage. Nearly all AIF respondents (369/398) received both textual and numeric feedback, whereas under two-thirds of PF students received no textual peer feedback at all; many peers delivered non-targeted praise rather than task-focused, improvement-oriented guidance. When high-quality textual peer feedback was received, PF performance matched AIF, indicating the AI's edge stems from consistent, reliable provision rather than inherent superiority.
-
AI feedback is valued less but activates more. Students rated peer feedback slightly higher on perceived validity and emotional response (evidence of mild algorithm aversion), yet the reliable provision of AI feedback still translated into higher participation and learning gains — perceived preference did not align with behavioral outcomes.
-
Personalization without sacrificing consistency. AIF offered individually tailored feedback at scale while maintaining uniform quality, overcoming the personalization–consistency trade-off that constrains peer and adaptive systems, suggesting GPT-4 can complement lecturer feedback and substitute for unreliable peer feedback in large classes.
Connected Concepts
Connected Articles
- AI Generated Feedback Higher Ed — AI-generated feedback in higher education
- Becerra Aicofe Feedback 2026 — AI-powered collaborative feedback
- Coach Not Crutch AI Writing — AI vs. human feedback on writing practice
- AI Vs Human Assessment Efl Tpck 2026 — AI vs. human-developed assessment tasks
- LLM Formative Feedback Systematic Review 2026 — Systematic review of LLM formative feedback
- GenAI Educational Outcomes Meta Analysis — Meta-analysis of generative AI learning outcomes
Citation
Geschwind, S., Graf Lambsdorff, J., Voss, D., & Hackl, V. (2026). GPT-4 feedback increases student activation and learning outcomes in higher education. International Journal of Artificial Intelligence in Education, 36, 100014.