Artificial intelligence and feedback in university education: effectiveness and student perceptions

Created: 2026-07-19 | Tags: generative-aifeedback-loophigher-edstudent-experiencelearning-gainsformative-assessmentai-literacy

Valentina Grion (Pegaso Telematic University), Beatrice Doria (Pegaso Telematic University), Daniele Agostini (University of Trento), Giorgia Slaviero (University of Padua) (2026)Assessment & Evaluation in Higher Education (Taylor & Francis). Open Access, CC BY 4.0. doi:10.1080/02602938.2026.2697962.

📄 Full text (Taylor & Francis, OA) — open supplementary materials (rubric, assignment, AI prompt protocol, anonymised datasets) at Zenodo 20814177

Summary

This quasi-experimental study directly compares AI-generated feedback (two LLMs: GPT-o4-mini and DeepSeek R1) with expert human-teacher feedback in a project-based university course (Assessment & Learning, third-year Primary Teacher Education, University of Padua). The central question is not "is AI feedback worse?" but under what pedagogical conditions AI feedback can be a credible, educationally meaningful component of formative assessment. The answer the authors land on: feedback effectiveness depends less on its source than on the pedagogical architecture in which it is embedded — strong assessment literacy and explicit, shared criteria make AI feedback comparable to teacher feedback.

Method (key parameters)

Key Findings

RQ1 — Feedback improves performance regardless of source

Across all 47 groups, project performance rose significantly from PRE to POST (Wilcoxon W = 1081, p < 0.001, rank-biserial rrb = 0.77 — a large effect); mean score +3.9 points (23.81 → 27.70), with post-test scores converging near the ceiling (median 28).^[raw/papers/tandf-2026-ai-generated-feedback-higher-ed.md]

RQ2 — No significant difference between feedback sources

Post-feedback scores did not differ by source (Kruskal–Wallis H(2) = 1.91, p = 0.384, ε² = 0.042); gain scores likewise non-significant (H(2) = 0.74, p = 0.690). Pairwise Hodges–Lehmann contrasts all had CIs spanning zero.^[raw/papers/tandf-2026-ai-generated-feedback-higher-ed.md]

RQ3 — Attendance doesn't matter

Robust linear model: no main effect of attendance (F(1,41) = 1.52, p = 0.225), no source × attendance interaction (F(2,41) = 0.97, p = 0.389).^[raw/papers/tandf-2026-ai-generated-feedback-higher-ed.md]

RQ4 — AI feedback is practically comparable to teacher feedback

Comparison (AI − Teacher) Mean diff 90% CI Non-inferior? Equivalent?
GPT-o4-mini vs Teacher +0.23 [−0.46, 0.91] Yes Yes
DeepSeek R1 vs Teacher +0.56 [−0.05, 1.18] Yes No (upper bound exceeds +1)

Same pattern on baseline-adjusted gains (DIFF_ADJ). GPT-o4-mini met both non-inferiority and full equivalence; DeepSeek R1 met non-inferiority (practically comparable, but with more uncertainty).^[raw/papers/tandf-2026-ai-generated-feedback-higher-ed.md]

Student perceptions — equally positive across sources

Validated 19-item questionnaire (N = 200; scales: perceived mastery α = 0.81, positive emotions α = 0.85, negative emotions α = 0.73). Students were blind to feedback source. No significant differences across conditions on any scale:

AI-generated feedback was experienced as acceptable and supportive, comparable to teacher feedback.^[raw/papers/tandf-2026-ai-generated-feedback-higher-ed.md]

Interpretation: Source vs. Architecture

The authors' core argument: feedback works as a systemic, relational process, not a function of who (or what) produces it. In this study both AI and teacher feedback were anchored to the same explicit rubric and student co-constructed exemplar, which made criteria transparent and gave the AI an "interpretative anchor" usually tacit in human grading. It is the teacher's assessment literacy — encoded in the rubric and exemplar — that calibrated the AI, not the model alone. Thus generative AI is best seen as a support for teachers with strong assessment literacy (scaling timeliness/consistency) rather than an autonomous replacement. The study explicitly warns against over-reliance and unequal access, and calls for maintaining teacher oversight and students' critical engagement.^[raw/papers/tandf-2026-ai-generated-feedback-higher-ed.md]

Limitations (per authors)

Implications for the wiki

Related Pages

Citation

APA: Grion, V., Doria, B., Agostini, D., & Slaviero, G. (2026). Artificial intelligence and feedback in university education: effectiveness and student perceptions. Assessment & Evaluation in Higher Education. https://doi.org/10.1080/02602938.2026.2697962