Research Article
Assessing the Impact and Underlying Pathways of Sequenced AI Feedback on Student Learning
Synthesis: Sequenced AI feedback harms learning despite boosting engagement and positive perceptions.
Core Finding
In a randomized experiment with 199 participants, the authors compared two types of AI-generated feedback:
- Sequenced (layered): Encouragement → hints → correct answer, designed to promote learner autonomy
- Non-sequenced (direct): Full-solution feedback immediately
Contrary to design intuition, sequenced feedback led to significantly poorer learning performance. The finding reveals a critical disconnect between what students like and what actually helps them learn.
Mediation Pathways
Three causal pathways were tested via mediation analysis:
| Pathway | Mediator | Effect | Significant? |
|---|---|---|---|
| Affective | Perceived encouragement | Positive → better learning | ✓ |
| Behavioral | Tasks needing ≥3 submissions | Negative → worse learning | ✓ |
| Cognitive | Mental effort | Neutral | ✗ |
The positive affective pathway (students felt more encouraged) was completely counteracted by the negative behavioral pathway (students made more resubmissions). The net effect was significantly poorer learning outcomes.
Key Mechanisms
Why Sequenced Feedback Backfired
- The hint-before-answer structure inadvertently encouraged trial-and-error behavior rather than deep processing
- Students submitted more attempts per task, indicating they were "gaming" the hint system rather than engaging in genuine Problem Solving
- Higher mental effort was reported but did not translate to better learning — suggesting the effort was directed at navigating the feedback sequence rather than understanding the material
Why Direct Feedback Worked
- Immediate corrective information eliminated the temptation to guess
- Students processed the solution rather than iterating through hints
- Lower engagement scores but higher learning outcomes
Design Implications
This study challenges the prevailing intuition that more scaffolded, autonomy-supportive feedback is always better. Key takeaways for AI feedback system design:
- Engagement ≠ learning: User satisfaction and behavioral engagement are not reliable proxies for learning gains — designers must measure learning outcomes directly
- Limit resubmission loops: Systems should cap hint requests or require reflection between attempts to prevent trial-and-error gaming
- Strategic blending: Consider providing direct corrective feedback first, with optional encouragement and hints available on demand rather than as a mandatory sequence
- Cognitive load management: The higher mental effort induced by sequenced feedback did not aid learning — design should channel effort toward understanding rather than navigation
Connection to Existing Knowledge Base
This paper directly informs several threads in the knowledge base:
- Formative Assessment: Direct evidence about AI-generated feedback design — sequencing that feels supportive may undermine formative goals
- Scaffolding critical thinking with generative AI: Design principles for integrating large language models in higher: Vendrell & Johnston's eight design principles for Large Language Models (LLMs) Scaffolding — this study provides empirical evidence that poorly designed scaffolding can harm learning, reinforcing the need for "cognitive friction" design
- Prober.ai: Gated Inquiry-Based Feedback via LLM-Constrained Personas for Argumentative Writing: The inverted paradigm (AI asks questions, gates suggestions) offers an alternative to sequenced feedback that may avoid the resubmission trap
- Self-Regulated Learning: Sequenced feedback was intended to promote autonomy and SRL, but the behavioral data shows it had the opposite effect — a cautionary tale for SRL-aligned AI design
- Metacognition: The engagement-learning disconnect exemplifies the metacognitive calibration problem — students felt they were learning more with sequenced feedback when they were actually learning less
- AICoFe: Implementation and Deployment of an AI-Based Collaborative Feedback System for Higher Education: Multi-LLM collaborative feedback systems must consider feedback sequencing carefully to avoid the pitfalls identified here
- The Pedagogy of AI Mistakes: Fostering Higher-Order Thinking: Hosseini's work on deliberately leveraging AI errors connects to the finding that easy, encouraging feedback may be less pedagogically effective than direct correction
- Transfer of Learning: The learning outcome disparity between conditions raises transfer implications — do sequenced-feedback students retain less when the scaffolding is removed?
Methodological Strengths
- Randomized controlled design with 199 participants — causal claims are well-supported
- Mediation analysis identifies why the effect occurs, not just whether it occurs
- Multi-dimensional measurement: learning performance, behavioral engagement (submission patterns), cognitive engagement (mental effort), and affective perceptions
- Multi-institution collaboration (UNC, CMU, Pitt, HKU)
Open Questions
- Would results differ with longer exposure (multi-session vs. single-session study)?
- Does domain matter — would sequenced feedback work better for ill-defined problems than well-defined ones?
- Can the resubmission problem be solved by requiring reflection prompts between hint levels?
- Would a hybrid design (direct feedback + optional hints) preserve learning while maintaining positive affect?
What this means for practice
- Learners. Do not guess your way up a hint ladder. In the layered condition, students who accumulated three or more submissions on a task scored lower (indirect effect = −0.19, p = .03), a sign of trial-and-error rather than thinking.
- Learners. When you already know what to fix, direct corrective feedback is the safer bet: the layered group scored significantly lower on the post-test overall (B = −0.83, p = .02) despite feeling more encouraged.
- Instructors. Cap resubmissions per task, or require a brief written reflection between attempts. Frequent submissions strongly predicted worse post-test scores (B = −0.33, p = .001).
- Instructors. Never read encouragement or satisfaction ratings as evidence of learning. Layered feedback raised perceived encouragement and mental effort, yet the effort self-report did not predict performance (p = .77) and the net effect on learning was negative.
- Instructors. Default to direct feedback when the goal is short-term performance in online higher education, and reserve sequenced designs for courses where motivation and independence are the priority — the design the authors note retains value for the long term.
Limitations
- The final dataset is 199 U.S. college students recruited through Prolific (100 layered, 99 non-layered, mean age 32) after attention-check screening of 215 — a convenience sample of online learners.
- The study ran in a single course context and is a single-session design with no delayed post-test, so long-term retention and transfer were not measured.
- The learning-by-doing tasks were relatively low-level by the authors' own account and may not have engaged higher-order thinking.
- Process data came mainly from log traces; there was no direct measure of how learners attended to or processed the feedback, and the authors call for richer instruments such as eye-tracking.
Citation
Cao, J., Zhao, C. Q., Schunn, C., McLaughlin, E. A., Lin, J., & Koedinger, K. R. (2026). Assessing the Impact and Underlying Pathways of Sequenced AI Feedback on Student Learning.