On this page

Feedback — information provided to a learner about their performance or understanding that is intended to close the gap between current and desired performance. In AI in education, feedback has become a central and rapidly transforming theme: AI systems now generate, deliver, and even teach students how to use feedback, reshaping every stage of the feedback process.

Questions to Consider

  • Feedback is often imagined as information handed to a learner, but research argues it only 'counts' as feedback when the learner makes sense of and acts on it. If feedback is a relationship rather than a message, what changes about how we should deliver it?
  • Students' perception of whether feedback came from an AI or a human can change how much they learn from it. Why do you think the source matters so much — and what does that imply for AI-generated feedback in your context?
  • Immediate feedback can be a trap: the 'correct-answer trap' shows that students who simply copy corrections learn less. When does instant feedback help, and when does it short-circuit the very learning it's meant to support?
  • Research even found that a well-designed sequence of feedback — encouragement, then hints, then the answer — designed to build autonomy actually harmed learning despite boosting engagement. If students feel good about feedback that makes them learn less, what should we optimize?
  • Feedback is described as a system with parts: the feedback loop, quality of provision, the learner's uptake capacity, and the assessment context. If you were trying to improve feedback in your own course, which part would you fix first, and why?

Introduction

This is the umbrella concept for the knowledge base's feedback-related ideas. Feedback sits at the intersection of assessment and learning: without feedback, assessment measures performance but does not improve it; with effective feedback, assessment becomes a learning event. The knowledge base treats feedback as a system with multiple facets — the quality of the feedback itself (AI Feedback Quality), the loop through which it closes the learning gap, the learner's capacity to use it (Feedback Literacy), and the assessment contexts in which it operates (Formative Assessment, Peer Review, Automated Assessment).

  • Yilmaz et al. show that students' perception of the feedback source (AI vs human) shapes self-regulated learning from GenAI feedback — a crucial qualifier for feedback effectiveness claims.

The feedback system

Feedback is best understood not as a single event but as a connected system of interacting parts, each of which the knowledge base documents as its own concept:

  • The feedback loop — the mechanism through which feedback closes the gap between current and desired performance; the cycle of performance → feedback → revision → improved performance that drives learning.
  • AI Feedback Quality — the accuracy, usefulness, timeliness, and pedagogical value of feedback generated by AI; the provision side of the system.
  • Feedback Literacy — the capabilities and dispositions students need to understand, evaluate, and act on feedback; the uptake side of the system.
  • Formative Assessment — the assessment context in which feedback is used to improve learning while it is still in progress, rather than merely to judge it.
  • Peer Review — feedback exchanged between learners, increasingly augmented by AI and a site for developing feedback literacy.
  • Automated Assessment — AI-driven scoring and feedback, from automated essay scoring to confidence-aware short-answer grading.

The feedback loop

The feedback loop is the cyclical process where AI systems assess student work, deliver feedback, observe the student's response, and adapt subsequent instruction. AI-mediated feedback loops operate at multiple timescales:

  • Immediate feedback: Automated grading systems and AI tutors provide real-time correction during Problem Solving. Correct-answer trap research shows that immediate feedback can short-circuit learning if students simply copy corrections. Generating that immediate corrective feedback with LLMs inside a tutor is feasible but imperfect: across 6,926 logged transactions in the Apprentice Tutor College Algebra platform, Reddig, Arora & MacLellan (2025) found GPT-4 produced hints targeted to the student's specific error ~66% of the time, yet ~35% of hints were too general, incorrect, or prematurely gave away the answer, and LLM-based automated quality checks misaligned with human judgment — reinforcing that unvetted AI feedback in the immediate loop carries real short-circuiting risk.
  • Assignment-level feedback: Formative Assessment systems and AI feedback quality research examine whether AI-generated assignment feedback improves subsequent work. Sequenced feedback studies test whether the order of feedback matters.
  • Course-level loops: Learning analytics dashboards and educational platforms aggregate feedback across assignments to identify patterns and recommend interventions.

The effectiveness of a feedback loop depends on feedback quality — accuracy, specificity, timeliness, and actionability. AI peer feedback systems add a social dimension to the loop.

The human in the loop: feedback loops are not purely automated — teachers often mediate AI-generated feedback before it reaches learners. Studies of AI feedback tools for teachers (e.g., the PolyFeed tool combining an ML detector with an LLM rephraser) find teachers use professional judgement to accept, edit, or reject AI suggestions — an "assist but verify" pattern — and systematically moderate exaggerated praise and generic suggestions to protect authenticity and voice. The relational/affective dimension of feedback (student–teacher relationship, encouragement) most strongly resists AI delegation, suggesting this part of the loop remains inherently human. This human-in-the-loop mediation connects to Human In The Loop AI and to Teacher Role.

How AI transforms feedback

AI changes feedback in three consequential directions, each raising the stakes of the other facets:

  • Volume and immediacy. AI can generate feedback instantly and at scale, dramatically increasing how much feedback students receive. Multi-site GenAI feedback studies and AI feedback in higher education document AI feedback experienced as comparable to teacher feedback in acceptability and supportiveness. A PRISMA-guided systematic review of 42 empirical studies (2023–2025) corroborates this scale benefit — LLMs can accelerate grading and deliver rapid, personalized feedback at scale, especially in large or higher-education cohorts — while cautioning that such feedback is sometimes too generic or misaligned with the assigned grade, and that reliability slips on longer, multilingual, or nuanced tasks (Jukiewicz Chatgpt Teacher Assessment Feedback 2026).
  • Evaluation demands on the learner. Because AI feedback can be inaccurate or hallucinated, students must judge whether to trust and act on it. This is where Feedback Literacy and Trust Calibration become decisive: Mendoza et al. (2026) show that only feedback-literate students convert AI feedback into Self Regulated Learning gains, while the LLM fallacy captures how students may over-credit AI feedback to their own competence.
  • Role-aware feedback. Yaşar et al. (2026) showed that prompting the same LLM to evaluate the same student artifact under different roles — instructor, peer reviewer, grant reviewer — produced qualitatively distinct feedback: instructors were encouraging and process-oriented, peers supportive and conversational, grant reviewers formal and outcomes-oriented. These differences were epistemic, not merely stylistic, foregrounding different aspects of design practice — evidence that role-aware prompting can generate role-sensitive evaluative feedback rather than a single generic response.
  • **Feedback literacy as a training goal.**single generic response.
  • Feedback utility vs. evaluative authority. AlGhamdi (2026) shows students can simultaneously accept AI-generated feedback as useful and reject AI as the appropriate grader — separating feedback utility (the informational value of surface-level feedback) from evaluative authority (who may decide the grade). In a transparent post-assessment design, Saudi computing students treated these as analytically distinct judgements, valuing ChatGPT's clarity for revision while reserving grading authority for the human instructor and articulating the dialogical, institutionally-weighted reasons why. The finding extends Feedback Literacy beyond judging feedback quality to reasoning about evaluative authority itself, and suggests transparency about AI involvement can make feedback an object of critical reflection rather than passive acceptance.
  • Feedback literacy as a training goal. Rather than only providing feedback, AI tools are increasingly designed to teach students to use feedback — reframing automated feedback tools as literacy-building instruments, framing GenAI as an enabler of feedback engagement, and using GenAI critique tasks to build psychological, feedback, and AI literacies together.

What makes feedback effective: calibration evidence

Ngai & Gilbert (2026) provide a clean experimental result on feedback design with direct relevance to AIED: veridical, immediate, trial-by-trial feedback tied to a learner's own prior prediction is what changes behavior — prediction or beliefs alone are not enough. In their four-group design, feedback combined with a preceding performance prediction improved calibration and behavior, while predictions without feedback did nothing, and adding an explicit over-/underconfidence label added nothing further. LLMs can also serve as calibration partners in this process: Yaşar et al. (2026) found that after rubric refinement, LLMs demonstrated greater consistency than some human raters in applying performance thresholds — useful for norming sessions and formative peer feedback environments, where the model helps calibrate evaluative judgment rather than replace it. This aligns with the knowledge base's feedback-system view: feedback's power lies in closing the gap against a learner's own estimate (the provision–uptake pairing), and effective feedback should be immediate, accurate, and explicitly connected to what the learner predicted — a design principle for AI feedback systems and Feedback Literacy training alike.

Feedback across assessment contexts

The knowledge base's feedback research spans the full range of assessment contexts, each with distinct feedback dynamics:

  • Formative feedback is the canonical site of feedback-for-learning — formative assessment feeds self-regulated learning through feedback that arrives while learning is still in progress (automated formative assessments, feedback enactment).
  • Summative feedback is the feedback attached to summative assessment — end-of-unit tests, oral exams, and proctored/closed-book examinations. In the AI era, the summative setting is where AI resistance matters most (see Summative Assessment): feedback on a proctored oral exam or closed-book assessment tests genuine learning, whereas feedback on AI-assisted homework can be inflated. Oral assessments reframe the feedback moment as a live, interactive dialogue — feedback becomes immediate, conversational, and inseparable from the assessment itself, which is precisely why they resist AI substitution.
  • Peer feedback adds a social layer — peer feedback develops feedback literacy and is increasingly AI-augmented (AI peer feedback, GenAI in EFL peer feedback). Yet peer feedback is only as good as its actual delivery: in a semester-long field experiment, under two-thirds of students assigned peer feedback received no textual feedback at all, and much of what was delivered was non-targeted praise — so when individual GPT-4 feedback replaced unreliable peers, students sustained the highest participation and posted the largest content learning gains (Geschwind et al., 2026), evidence that the AI's edge is a reliability advantage rather than inherent superiority.
  • Automated feedback scales delivery — automated scoring and feedback from essay scoring to short-answer grading. Arthur (Yin et al. 2026) extends this to unstructured quantitative problems: a per-question XGBoost diagnosis backbone infers rubric-labelled mistakes from students' submitted numerical answers to Engineering Economics Calculated Formula Questions and turns them into natural-language feedback, while a dialogue-based scheme requests intermediate answers only when confidence is low — balancing feedback accuracy against the collection burden on students.

A key cross-context insight is that the reliability of feedback depends on the integrity of the assessment it is attached to: feedback is only as trustworthy as the measure it responds to. In the AI era this pushes educators toward authentic and AI-resistant summative formats where the feedback a student receives reflects genuine learning rather than AI-assisted output.

The "assessment for learning" paradigm reframes feedback as an overarching philosophy rather than a component. Mesny, Roberge-Maltais & Galy (2026) synthesize the wider higher-education literature to frame formative, ongoing, and individualized feedback as central to that philosophy, balancing formative with summative purposes and foregrounding student agency, self-Regulation, and metacognitive skill. They note, however, that in management education feedback-rich practices remain unevenly adopted: self- and peer-assessment (a site of peer feedback) dominate the literature, while reassessment — which uses feedback-driven second chances to improve learning — is virtually absent.

The provision-uptake pairing

The knowledge base's core feedback insight is that feedback quality and feedback literacy are two sides of one system: high-quality feedback is inert without a literate recipient, and a literate student gains little from poor feedback. AI Feedback Quality covers the provision side (is the feedback accurate, timely, actionable?), while Feedback Literacy covers the uptake side (can the student judge and act on it?). The feedback loop is what connects them — the mechanism by which quality feedback, received by a literate learner, closes the gap. Designing effective AI feedback therefore means designing both the system and the student.

A two-layer model shows how the pairing can be organized around a machine without delegating judgement to it. In El Khoury and Ma's worked example, an agent produces a draft rubric-based evidence report on each transcript with every rating tied to quoted excerpts, and the instructor then reviews the report, leads a debrief and decides what the evidence means for that student — a division the authors summarise as the AI organizing evidence while the instructor interprets it. Their claim about uptake is appraisal-based: feedback and iteration outside the social hierarchies students navigate with peers and instructors are easier to attempt, and rehearsal at the student's own pace is what they argue turns occasional confidence into a settled habit of engaging with feedback.

Why feedback matters for AI in education

Feedback is one of the most consequential and best-evidenced mechanisms in education, and AI both amplifies and complicates it. Well-architected AI feedback can match or exceed human feedback and scale across cohorts, but it demands new learner capabilities (Feedback Literacy, AI Literacy) and carries risks (uncritical acceptance, Over-Reliance). As AI-generated feedback becomes ubiquitous, the knowledge base frames feedback as a whole system — quality, loop, literacy, and assessment context working together — rather than as any single component.

Connected Concepts

Connected Articles

Connected FAQs

Embed this page

Copy the code below to embed a chromeless version of this page in a learning management system or other website. The embedded view hides the site header, navigation, and footer.