On this page

Synthesis: Sequenced AI feedback harms learning despite boosting engagement and positive perceptions.

Core Finding

In a randomized experiment with 199 participants, the authors compared two types of AI-generated feedback:

  • Sequenced (layered): Encouragement → hints → correct answer, designed to promote learner autonomy
  • Non-sequenced (direct): Full-solution feedback immediately

Contrary to design intuition, sequenced feedback led to significantly poorer learning performance. The finding reveals a critical disconnect between what students like and what actually helps them learn.

Mediation Pathways

Three causal pathways were tested via mediation analysis:

Pathway Mediator Effect Significant?
Affective Perceived encouragement Positive → better learning ✓
Behavioral Tasks needing ≥3 submissions Negative → worse learning ✓
Cognitive Mental effort Neutral ✗

The positive affective pathway (students felt more encouraged) was completely counteracted by the negative behavioral pathway (students made more resubmissions). The net effect was significantly poorer learning outcomes.

Key Mechanisms

Why Sequenced Feedback Backfired

  • The hint-before-answer structure inadvertently encouraged trial-and-error behavior rather than deep processing
  • Students submitted more attempts per task, indicating they were "gaming" the hint system rather than engaging in genuine Problem Solving
  • Higher mental effort was reported but did not translate to better learning — suggesting the effort was directed at navigating the feedback sequence rather than understanding the material

Why Direct Feedback Worked

  • Immediate corrective information eliminated the temptation to guess
  • Students processed the solution rather than iterating through hints
  • Lower engagement scores but higher learning outcomes

Design Implications

This study challenges the prevailing intuition that more scaffolded, autonomy-supportive feedback is always better. Key takeaways for AI feedback system design:

  1. Engagement ≠ learning: User satisfaction and behavioral engagement are not reliable proxies for learning gains — designers must measure learning outcomes directly
  2. Limit resubmission loops: Systems should cap hint requests or require reflection between attempts to prevent trial-and-error gaming
  3. Strategic blending: Consider providing direct corrective feedback first, with optional encouragement and hints available on demand rather than as a mandatory sequence
  4. Cognitive load management: The higher mental effort induced by sequenced feedback did not aid learning — design should channel effort toward understanding rather than navigation

Connection to Existing Knowledge Base

This paper directly informs several threads in the knowledge base:

Methodological Strengths

  • Randomized controlled design with 199 participants — causal claims are well-supported
  • Mediation analysis identifies why the effect occurs, not just whether it occurs
  • Multi-dimensional measurement: learning performance, behavioral engagement (submission patterns), cognitive engagement (mental effort), and affective perceptions
  • Multi-institution collaboration (UNC, CMU, Pitt, HKU)

Open Questions

  • Would results differ with longer exposure (multi-session vs. single-session study)?
  • Does domain matter — would sequenced feedback work better for ill-defined problems than well-defined ones?
  • Can the resubmission problem be solved by requiring reflection prompts between hint levels?
  • Would a hybrid design (direct feedback + optional hints) preserve learning while maintaining positive affect?

What this means for practice

  • Learners. Do not guess your way up a hint ladder. In the layered condition, students who accumulated three or more submissions on a task scored lower (indirect effect = −0.19, p = .03), a sign of trial-and-error rather than thinking.
  • Learners. When you already know what to fix, direct corrective feedback is the safer bet: the layered group scored significantly lower on the post-test overall (B = −0.83, p = .02) despite feeling more encouraged.
  • Instructors. Cap resubmissions per task, or require a brief written reflection between attempts. Frequent submissions strongly predicted worse post-test scores (B = −0.33, p = .001).
  • Instructors. Never read encouragement or satisfaction ratings as evidence of learning. Layered feedback raised perceived encouragement and mental effort, yet the effort self-report did not predict performance (p = .77) and the net effect on learning was negative.
  • Instructors. Default to direct feedback when the goal is short-term performance in online higher education, and reserve sequenced designs for courses where motivation and independence are the priority — the design the authors note retains value for the long term.

Limitations

  • The final dataset is 199 U.S. college students recruited through Prolific (100 layered, 99 non-layered, mean age 32) after attention-check screening of 215 — a convenience sample of online learners.
  • The study ran in a single course context and is a single-session design with no delayed post-test, so long-term retention and transfer were not measured.
  • The learning-by-doing tasks were relatively low-level by the authors' own account and may not have engaged higher-order thinking.
  • Process data came mainly from log traces; there was no direct measure of how learners attended to or processed the feedback, and the authors call for richer instruments such as eye-tracking.

Citation

Cao, J., Zhao, C. Q., Schunn, C., McLaughlin, E. A., Lin, J., & Koedinger, K. R. (2026). Assessing the Impact and Underlying Pathways of Sequenced AI Feedback on Student Learning.

Embed this page

Copy the code below to embed a chromeless version of this page in a learning management system or other website. The embedded view hides the site header, navigation, and footer.