Research Article
Human-centered GenAI feedback design in higher education: a multisite experiment on direct, reflective, and hybrid approaches to scientific argumentation
Synthesis: A multisite, cluster-randomized field experiment (1,176 first-year undergraduates, 48 sections, 4 universities, 3 science domains) compares four Feedback designs for scientific argumentation: peer-only, direct GenAI, reflective GenAI (self-evaluation then AI critique), and hybrid (self-evaluation + peer + GenAI). The hybrid condition produced the highest argument-quality gains and clearest advantage on conceptual learning; reflective and hybrid designs both outperformed direct GenAI on delayed AI-free transfer. Findings suggest that GenAI's educational value depends less on AI access than on preserving student Learner Agency, evaluative judgment, and ownership during revision.
Study Design
Ateş conducted a multisite, cluster-randomized, longitudinal field experiment in introductory university science courses:
- 1,176 first-year undergraduates from 48 course sections across 4 universities
- 3 science domains — biology, chemistry, physics
- 4 feedback conditions randomized at the section level:
- Peer feedback only (control)
- Direct GenAI-supported feedback — AI critique delivered to students
- Reflective GenAI-supported feedback — self-evaluation first, then AI critique
- Hybrid design — self-evaluation → peer feedback → GenAI critique
Key Findings
| Outcome | Direct GenAI | Reflective GenAI | Hybrid |
|---|---|---|---|
| Immediate argument-quality gain | Better than peer | — | Highest |
| Feedback uptake | — | Stronger | Stronger |
| Self-regulated learning | — | Stronger | Stronger |
| Conceptual learning | — | Positive (n.s.) | Clearest advantage |
| Delayed AI-free transfer | — | Outperformed direct | Outperformed direct |
- Direct GenAI improved immediate argument quality over peer feedback but showed weaker transfer
- Reflective and hybrid designs produced stronger feedback uptake and self-regulated learning
- Hybrid condition showed the clearest advantage on conceptual learning
- Both reflective and hybrid outperformed direct on delayed AI-free transfer
- Multilevel mediation: feedback uptake and self-regulated learning partially explained these advantages
Why Design Matters
The paper argues that feedback becomes educationally valuable not through comment delivery alone, but when learners:
- Interpret critique
- Compare it against criteria
- Judge its relevance
- Use it to improve subsequent work
Direct GenAI Feedback may encourage passive uptake — students outsource evaluative judgment to the system. Reflective and hybrid designs preserve epistemic agency: the student must first evaluate their own work, compare peer/AI inputs, and decide how to revise.
The core insight: GenAI's educational value depends less on AI access per se than on whether feedback environments preserve student agency, evaluative judgment, and ownership during revision.
g revision.**
What this means for practice
- Instructors. Sequence feedback so students evaluate their own work before reading AI critique: the reflective and hybrid conditions produced stronger feedback uptake, stronger self-regulated learning, and better delayed AI-free transfer than direct AI feedback.
- Instructors. Pair AI critique with peer feedback instead of substituting it — the hybrid condition (self-evaluation → peer feedback → GenAI critique) showed the highest immediate argument-quality gains and the clearest advantage on conceptual learning.
- Instructional designers. Build the revision process into the task, not just the feedback channel: gains were partially mediated by feedback uptake and self-regulated learning during revision, and the hybrid revision memo was capped at 100–150 words to force judgment rather than transcription.
- Faculty developers. Model and require the four revision steps the paper identifies — interpret critique, compare it against criteria, judge its relevance, then revise — since direct GenAI encourages students to outsource evaluative judgment.
- Administrators. Fund cross-site comparability work: a shared task-design protocol and a common analytic rubric across 4 universities and 3 science domains let the 48 sections be compared at all.
Limitations
- Of 1,248 enrolled students, 72 were excluded for non-consent, course withdrawal before the first cycle, or absence from both the post-test and delayed-transfer sessions, and because research consent was voluntary the analytic sample may carry consent-based selection.
- Self-regulated learning rested substantially on self-report (a 12-item task-specific scale, McDonald's ω = .88) blended with LMS trace indicators, and feedback uptake was a four-indicator construct.
- Argument-quality gains faced ceiling-related constraints on dimensions where drafts were already strong, which the authors addressed only through a supplementary baseline-adjusted final-score sensitivity model.
- The 48-section cluster-randomized design assumed an intraclass correlation of .05 and was powered to detect effects of d = .25 and above, so smaller differences between conditions are not resolvable.
Citation
Ateş, H. (2026). Human-centered GenAI feedback design in higher education: A multisite experiment on direct, reflective, and hybrid approaches to scientific argumentation. International Journal of Educational Technology in Higher Education, 23(38)