Summary
This quasi-experimental study directly compares AI-generated feedback (two LLMs: GPT-o4-mini and DeepSeek R1) with expert human-teacher feedback in a project-based university course (Assessment & Learning, third-year Primary Teacher Education, University of Padua). The central question is not "is AI feedback worse?" but under what pedagogical conditions AI feedback can be a credible, educationally meaningful component of formative assessment. The answer the authors land on: feedback effectiveness depends less on its source than on the pedagogical architecture in which it is embedded β strong assessment literacy and explicit, shared criteria make AI feedback comparable to teacher feedback.
Method (key parameters)
Design: 47 student groups (N = 238; 146 attending, 92 non-attending) randomly assigned to one of three feedback conditions β DeepSeek R1 (16 groups), expert human teacher (16), GPT-o4-mini (15). Unit of analysis = group (4β5 students each) to preserve independence.Task: Two-stage project (PRE then POST), evaluated with a shared analytic rubric (0β30 points) co-constructed with students.AI prompt design: Both LLMs were given all course materials plus an assignment brief, the pedagogical framework, and the co-constructed rubric via a Retrieval-Augmented Generation (RAG) setup; instructed to act as a university professor and give objective, justified, actionable formative feedback. The rubric + an exemplar functioned as a "calibration device" that transferred the teacher's evaluative expectations into the AI.Analyses: Wilcoxon signed-rank (PREβPOST), KruskalβWallis across sources, robust linear models (HC3) for attendance moderation, and β crucially β non-inferiority and equivalence tests (Welch-adjusted 90% CIs, pre-specified margin Β±1 point on the 30-point scale), because non-significant differences don't imply practical equivalence.Key Findings
RQ1 β Feedback improves performance regardless of source
Across all 47 groups, project performance rose significantly from PRE to POST (Wilcoxon W = 1081, p < 0.001, rank-biserial rrb = 0.77 β a large effect); mean score +3.9 points (23.81 β 27.70), with post-test scores converging near the ceiling (median 28).
RQ2 β No significant difference between feedback sources
Post-feedback scores did not differ by source (KruskalβWallis H(2) = 1.91, p = 0.384, Ρ² = 0.042); gain scores likewise non-significant (H(2) = 0.74, p = 0.690). Pairwise HodgesβLehmann contrasts all had CIs spanning zero.
RQ3 β Attendance doesn't matter
Robust linear model: no main effect of attendance (F(1,41) = 1.52, p = 0.225), no source Γ attendance interaction (F(2,41) = 0.97, p = 0.389).
RQ4 β AI feedback is practically comparable to teacher feedback
| Comparison (AI β Teacher) | Mean diff | 90% CI | Non-inferior? | Equivalent? |
|---|
| GPT-o4-mini vs Teacher | +0.23 | [β0.46, 0.91] | Yes | Yes |
| DeepSeek R1 vs Teacher | +0.56 | [β0.05, 1.18] | Yes | No (upper bound exceeds +1) |
Same pattern on baseline-adjusted gains (DIFF_ADJ). GPT-o4-mini met both non-inferiority and full equivalence; DeepSeek R1 met non-inferiority (practically comparable, but with more uncertainty).
Student perceptions β equally positive across sources
Validated 19-item questionnaire (N = 200; scales: perceived mastery Ξ± = 0.81, positive emotions Ξ± = 0.85, negative emotions Ξ± = 0.73). Students were blind to feedback source. No significant differences across conditions on any scale:
Perceived mastery: M β 4.14β4.22 (high)Positive emotions: M β 3.99β4.21Negative emotions: M β 1.22β1.39 (low)Overall satisfaction: ~98% (DeepSeek 97.5%, teacher 94%, GPT-o4-mini 100%) β analysed descriptively due to ceiling.AI-generated feedback was experienced as acceptable and supportive, comparable to teacher feedback.
Interpretation: Source vs. Architecture
The authors' core argument: feedback works as a systemic, relational process, not a function of who (or what) produces it. In this study both AI and teacher feedback were anchored to the same explicit rubric and student co-constructed exemplar, which made criteria transparent and gave the AI an "interpretative anchor" usually tacit in human grading. It is the teacher's assessment literacy β encoded in the rubric and exemplar β that calibrated the AI, not the model alone. Thus generative AI is best seen as a support for teachers with strong assessment literacy (scaling timeliness/consistency) rather than an autonomous replacement. The study explicitly warns against over-reliance and unequal access, and calls for maintaining teacher oversight and students' critical engagement.
Limitations (per authors)
Ceiling effect: 91% of groups scored β₯27/30 (SD = 0.95) β limits sensitivity of post-test comparisons; equivalence rests mainly on adjusted-gain analyses.Small group-level N = 47 β wide CIs; modest source differences can't be fully ruled out.No prior-AI-experience data collected; single course / discipline (Primary Teacher Education); student assessment literacy not measured (treated as a hypothesis, not tested).Implications for the wiki
A strong, well-controlled (randomised, blind, equivalence-tested) data point that AI-generated feedback can match expert teacher feedback for project-based learning when criteria are explicit and assessment literacy is high β complementing AI Feedback Quality and AI Learning Companions Framework work.Pairs naturally with Generative AI Guardrails Harm Learning (the PNAS RCT): that study shows unguarded AI tutoring can harm learning, this one shows well-architected AI feedback can match teachers β together they bracket the design-dependence of AIED outcomes.Reinforces Formative Assessment, Feedback Loop, and AI Literacy (teacher and student) as the decisive variables, over the raw tool.Connects to retrieval-augmented generation as a calibration mechanism and to Over Reliance (the authors flag it as a risk even in a positive-result study).Connected Concepts
AI LiteracyFormative AssessmentHigher EdRAGScaffoldingStudent ExperienceGenerative AILLMConnected Articles
AI Learning Companions Framework β Building AI Companions that Prioritise Learning over PerformanceCare Full Feedback GenAI β The care-full craft of feedback in an age of generative AIGenAI Teacher Feedback Comparison β Comparing Generative AI and teacher feedback: student perceptions of usefulness and trustworthinessGenerative AI Guardrails Harm Learning β Generative AI without guardrails can harm learning: Evidence from high school mathematicsLearner Centered Feedback AI β Enhancing learner-centered feedback with AI: teachers' practices and perceptionsRepeated AI Writing Feedback Semester β Student Evaluation of Repeated AI Feedback Across a Semester of WritingA4l Analytics Pipeline β Generalizing a Highly Configurable Analytics Pipeline to Replicate and Support Educational Research Across Multiple D...Aaai2026 Prompting Literacy K12 β Learning to Use AI for Learning: Teaching Responsible Use of AI Chatbot to K-12 Students Through an AI Literacy ModuleAcademiclaw Student Agent Benchmark β AcademiClaw: When Students Set Challenges for AI AgentsAccess Not Enough AI Tutoring 2026 β Access is Not Enough: Human Support Improves Engagement with AI TutoringAdapt Adaptive Lesson Plan Transformer β AdaPT: Adaptive Lesson Plan Transformer for Cross-Regional and Differentiated InstructionAdaptive Pretesting Retention β Do Gains from Generative AI-Enabled Adaptive Pretesting Persist? Evidence from a Retention StudyAffective Text Wearable Student Health β A Formative Study of Brief Affective Text as a Complement to Wearable Sensing for Longitudinal Student Health MonitoringAgency Gap AI Writing β The agency gap in AI-supported writing: how reactive and proactive agent designs shape multimodal reasoningAgent Voice Accents K12 Group Learning β Exploring How Agent Voice Accents Shape Human-AI Collaboration in K-12 Group LearningAgentic AI Education Scoping Review β Agentic AI in Education: A Scoping Review of Research Landscape, Capabilities, and the Frontier Agent ParadigmAgentic AI Pedagogical Best Practice 2026 β Agentic AI and Pedagogical Best Practice: The Tension Between Automation and LearningAgentic Education Coding β Agentic Education with AI Coding AssistantsAgentic Literacy Debt β Agentic Literacy Debt: A Structural Problem the AI Literacy Field Has Not Yet NamedAgents That Teach Incidental Learning β Agents That Teach: Designing Incidental Learning Back into AI-Assisted Software DevelopmentAgreement Not Quality LLM Coding Verification β Agreement Is Not Quality: Blind Expert Verification of Human and LLM Qualitative Coding When Human Consensus Is Not G...AI Adoption Training Public Sector β The Main Barrier to AI Adoption in the Public Sector is Lack of TrainingAI Adult Learning Guidelines Dis2026 β Guidelines for Designing AI Technologies to Support Adult LearningAI Agents Constructive Conflict Design Education 2026 β Enacting Constructive Conflicts with AI Agents to Enhance Reconsideration among Novice Interaction DesignersAI Agents Peer Learning Discourse β When AI Agents Teach Each Other: Discourse Patterns Resembling Peer Learning in the Moltbook CommunityCitation
Grion, V., Doria, B., Agostini, D., & Slaviero, G. (2026). Artificial intelligence and feedback in university education: effectiveness and student perceptions. Assessment & Evaluation in Higher Education. https://doi.org/10.1080/02602938.2026.2697962