On this page

Formative assessment — assessment designed to inform ongoing instruction and learning, as opposed to summative evaluation. In AI education, formative assessment is both transformed by AI and essential to it: AI systems can generate, validate, and adapt formative items and feedback at scale, while formative feedback is a primary mechanism through which AI tutors and adaptive systems support learning. The knowledge base's research examines AI-generated formative items, AI-generated feedback, and the design and evaluation of these systems.

Questions to Consider

  • Formative assessment is meant to close the loop — surface what students don't know and give feedback they can act on while learning is still in progress. How is that fundamentally different from summative evaluation, and when might the two get confused in practice?
  • AI can generate multiple-choice questions with impressive accuracy on verifiable dimensions, but is weakest on instructional-judgment dimensions. If a machine is good at correct answers but weaker at pedagogical judgment, what should it be trusted to do — and what should humans keep doing?
  • AI-generated feedback only helps when students actually enact it — the 'enacted feedback' condition, where students select, evaluate, and apply suggestions, outperformed simply being given feedback. If enactment matters more than the feedback itself, what does that mean for how formative feedback should be designed?
  • Feedback is described not as information transfer but as an ethical, relational practice. What gets lost when formative assessment is mass-produced as 'AI slop' — and what can human comment banks and relational care preserve?

Introduction

Formative assessment is central to AI in education because it sits at the junction of Assessment and learning feedback. Its purpose is to close the loop: surface what students know and don't know, and provide feedback they can act on to improve. AI makes this feasible at scale — generating items, scoring responses, and delivering individualized feedback — but the knowledge base's research shows that quality varies dramatically across item types, and that feedback only helps when students actually enact it.

AI-generated formative items

AI systems generate formative assessment items across modalities, with reliability varying by type:

AI-generated feedback

A large body of knowledge base research examines AI-generated formative feedback:

  • The enactment problem: Making AI-Generated Feedback Matter (13,037 students; 51,296 resources) shows feedback value depends on whether students enact it — the Enacted Feedback condition, where students select, evaluate, and apply AI feedback suggestions, outperformed simple directed feedback.
  • Feedback is not information transfer: The care-full craft of feedback argues feedback is an ethical, relational practice, not information transmission — feedback only constitutes feedback when students make sense of and act on it, and contrasts mass-produced "AI slop" with human comment-bank shortcuts.
  • Sequenced feedback can backfire: Sequenced AI feedback (encouragement → hints → correct answer, designed to promote autonomy) actually harmed learning in a randomized experiment (N=199) despite boosting engagement and positive perceptions — a cautionary finding about feedback design.
  • Learner-centered tools: PolyFeed combines ML suggestion models with teacher practice, showing how teachers adopt and adapt AI feedback suggestions; AI-supported internal feedback helps undergraduates develop evaluative judgment.
  • Rubric-guided prompting and role-aware feedback: Yaşar et al. (2026) showed that iterative rubric co-refinement — clarifying performance descriptors and explicitly accepting implicit indicators of learning — drove LLM–human agreement on student design work from 54.75% to 81.25% (Cronbach's Alpha rising from 0.393 to 0.798), with the largest gains in the cognitively demanding Iteration & Reflection category. Prompting the same model under instructor, peer-reviewer, and grant-reviewer roles produced distinct evaluative feedback, and post-revision LLMs were more consistent than some human raters in applying performance thresholds — positioning rubric-guided LLMs as calibration and co-design partners in formative feedback environments, with human-in-the-loop oversight remaining essential.
  • AI feedback sustains participation and drives gains at scale: Geschwind et al. (2026)'s semester-long field experiment across undergraduate tutorials found that individual GPT-4 formative feedback (spanning all three Hattie & Timperley dimensions — Feed-Back, Feed-Up, Feed-Forward) sustained the highest participation across eight open-ended tasks, lengthened student answers, and produced the strongest content learning gains — an effect driven by reliable, consistent AI provision, since when high-quality textual peer feedback was actually received, peer outcomes matched AI's.
  • Feedback futures: Feedback Futures synthesizes a special issue and argues the question is not whether GenAI can produce feedback but how to design feedback that supports learning, distilling recurring tensions across the field.
  • Diagnosis-first feedback for open-ended quantitative problems: Arthur (Yin et al. 2026) delivers real-time, personalized formative feedback on Engineering Economics Calculated Formula Questions, a domain where handwritten, unstructured solutions had previously blocked AI support. A per-question XGBoost backbone diagnoses likely rubric-labeled mistakes from students' submitted numerical answers (average precision 0.81, recall 0.79), and a dialogue-based scheme requests intermediate answers only when prediction confidence is low — balancing feedback accuracy against collection efficiency within a question-bank web interface.
  • Scale and limits of LLM formative feedback (systematic evidence): a PRISMA-guided systematic review of 42 empirical studies (2023–2025) finds LLMs can reduce teacher workload and deliver rapid, personalized feedback at scale — especially in large or higher-education cohorts — but that feedback is sometimes too generic or misaligned with the assigned grade and reliability slips on longer, multilingual, or nuanced tasks, reinforcing that formative AI feedback is best deployed under educator oversight (Can ChatGPT Replace the Teacher in Assessment? A Review of Research on the Use of Large Language Models in Grading and Providing Feedback).
  • Adaptivity is a separable ingredient, not decoration: Sukjaitham, Schaaf, Brod & Breitwieser (2026) supply the direct causal test that most LLM-feedback studies assume away, pitting GPT-4 response-contingent feedback against expert-written generic guidance matched on structure, tone, length, and motivational phrasing (verified with a five-dimension quality rubric, κ = .76–1.00). In a preregistered within-subjects experiment, 155 German fifth- and sixth-graders (M = 12.08 years) revised six if-then plans: plan quality rose from a median of 2 → 5 under adaptive feedback versus 2 → 3 under generic guidance (within-person V = 10,440, p < .001, r = .86; condition × time interaction estimate = 1.68, SE = 0.14, p < .001, with no pre-support difference). Children rated adaptive feedback both more helpful (r = .67) and more motivating (r = .74), and trial-level perceptions predicted the size of revision gains — making perceived usefulness part of the pathway rather than an affective byproduct. Because the control was itself well designed, the study shows added value beyond good non-contingent guidance rather than the difference between feedback and nothing: generic guidance is a genuine but limited substitute that plateaus at its own median. Planning served as the test case as a core self-regulated learning strategy with explicit quality criteria, which makes contingency — the property that distinguishes Scaffolding from static support — directly measurable on a one-sentence response. The authors define adaptivity narrowly as response-contingent adaptation of feedback content to the learner's concrete response, distinguishing it from conversational interactivity, tone, and stable-trait adaptive learning.
  • Perceived usefulness tracks actionability: Mendonça et al. (2026) had 144 programming students rate 893 LLM-generated feedback instances on five dimensions and 237 consolidated reports on six, holding domain, task, and instrument constant while educational level varied. Ratings were favorable throughout, with student-level means of 4.24 to 4.43 for individual answers and 4.11 to 4.38 for reports, yet actionability and usefulness drew the lowest ratings even as actionability and perceived accuracy carried the largest relative weights (32.6% and 29.0%) in a model explaining 76% of the variance in perceived usefulness, and motivation and personalization led a report-level model of intention to use explaining 53%. Since actionability is the dimension Feedback research treats as hardest to provide, the pattern reads as technology acceptance applied to feedback: usefulness tracks whether a learner can act, not how polished the feedback sounds.
  • Comprehension outruns actionability on the path to revision: Yao and Fan (2026) separated issue-level diagnosis from suggestions and gave students a written interpretation sheet (restatement, effect explanation, uncertainty identification, revision plan) before independent revision, structured teacher review, reassessment and reflection. Across three writing cycles in two intact classes (96 students; 288 cycle-level observations), feedback comprehension carried the strongest association with revision quality (r = 0.501, ahead of actionability at 0.480 and diagnostic accuracy at 0.445), and the diagnostic group gained 5.92 writing points against 3.58 (Hedges' g = 1.12), locating the leverage in students' understanding of the diagnosis rather than in its actionability alone.

Curriculum-grounded and educator-in-the-loop design

LearnLens addresses three persistent problems in AI formative assessment: error-aware assessment (capturing nuanced reasoning errors rather than surface mistakes), topic-linked memory chains (replacing noisy similarity-based RAG (Retrieval-Augmented Generation) with structured curriculum-grounded retrieval), and educator-in-the-loop design (teacher customization and oversight, not full automation). This connects to the broader tension in Human-in-the-Loop: scalable automation with expert validation.

Hoppe, Loibl & Leuders (2026) sharpen what educator-in-the-loop means once a tool produces its own claims rather than raw observations. Their conceptual analysis argues that AI-generated diagnostic inferences are qualitatively different evidence, because they are already the product of algorithmic interpretation, so teachers need a further layer the authors call meta-diagnosis: deliberately accepting, rejecting, or modifying an inference and integrating it with their own contextual knowledge. Specifying the DiaCoM framework for this treats AI-generated inferences as a situation characteristic and the accept, reject, or modify decision as diagnostic behavior, while expanding the person characteristics teachers need to include knowledge of how AI systems actually work. Because current systems rest mostly on performance data such as task correctness and completion time, motivational states and classroom dynamics stay largely absent, so the paper keeps teachers, not the dashboard, as the responsible reflective agents and frames judging algorithmic claims as a target for professional development.

Design trade-offs

Dimension AI Suitability Human Requirement
Factual correctness High Low
Concept alignment High Medium
Distractor quality Low High
Feedback depth Low High
Rubric consistency Medium Medium

Assessment, feedback, and learning

Formative assessment in AI education connects to the learning process itself:

  • Feedback loops: feedback loops are the mechanism by which formative assessment informs learning; AI tutors and adaptive systems close these loops at scale.
  • Self-regulated learning: formative feedback supports self-regulated learning when students monitor progress and adjust; AI feedback should cultivate evaluative judgment, not displace it.
  • Automated scoring reaches some phases of the cycle, not all: Chen & Liu (2026) ran a 14-week quasi-experiment in which 46 interpreting students submitted weekly renditions to an automated scoring system that returned an immediate score, transcript, marked errors, and a reference rendition, while a control group received whole-class teacher feedback only. The automated group improved more overall (d = 1.03), but the gain stayed where the deficit was decomposable, the signal reliable, and the scale sensitive: linguistic accuracy and logical coherence rose while information fidelity (agreement with human raters r = 0.12) and delivery fluency did not move. Self-regulated learning was uneven in a matching way, with execution and monitoring correlating with score gains (r = 0.42) while planning and emotional motivation sat near the scale midpoint, and several students deferred engagement after low scores instead of analyzing causes. The system reached the performance phase of the cycle, not the planning that precedes it.
  • From automated diagnosis to generated practice: Zhu, Luo & Li (2026) wire audio-score alignment, error quantification, and a Proximal Policy Optimization layer into one closed loop for music education, turning rhythm error signals (91.2% recall, 98.4% specificity) into rewards that steer generated practice tracks, and report a significant Group x Time interaction favoring the system group (beta = 0.52) across a 12 week quasi-experiment with 120 undergraduate music majors. The authors frame this as technical feasibility, and the practical reading is triage rather than assessment: precision of 89.7% means roughly one flagged rhythm error in ten is a false alarm, and the study's expert reviewers rated support for musical expression as the weakest area, so automated loops suit technical drills while expressive judgment stays with teachers.
  • Scaffolding: Scaffolding and formative assessment work together — AI can provide just-in-time hints and prompts, though sequenced feedback research cautions against over-structuring.
  • Validity and quality: the quality and validity of AI-generated formative items and feedback must be evaluated; AI Ed Evaluation provides the methods.

Risk: Assessment as surveillance

Formative assessment systems can shift from learning-support tools to behavior-monitoring infrastructure. The same data streams that enable adaptive tutoring can enable punitive tracking if AI Governance is weak. This connects to Privacy and student well-being, and argues for formative systems that support learning rather than surveil it.

Implications for AI in education

  • Match item type to AI reliability: use AI for verifiable dimensions (concept alignment, correctness) and retain human judgment for instructional dimensions (distractor quality, feedback depth).
  • Design for enactment, not just provision: AI feedback only helps when students select, evaluate, and apply it — structure workflows that support enactment.
  • Feedback design matters more than volume: sequenced or over-structured feedback can backfire; prioritize feedback that supports student sense-making and autonomy.
  • Keep educators in the loop: curriculum-grounded, educator-in-the-loop systems improve relevance and reduce noise.
  • Evaluate quality and validity: assess AI-generated items and feedback for quality, validity, and equity, not just generation speed.
  • Treat formative assessment as a philosophy, not a toolkit. Mesny, Roberge-Maltais & Galy (2026) synthesize the wider higher-education literature into an "assessment for learning" paradigm that frames formative, ongoing, and individualized Feedback as an overarching philosophy rather than a mere toolkit — balancing formative with summative purposes and foregrounding student agency, self-regulation, and metacognitive skill. They identify five mutually reinforcing practices (authentic assessment, self- and Peer Assessment, reassessment, standards-based grading, ungrading) through which this philosophy can be enacted, while noting their uptake remains highly uneven across higher-education fields.

Connected Concepts

Connected Articles

Connected FAQs

Embed this page

Copy the code below to embed a chromeless version of this page in a learning management system or other website. The embedded view hides the site header, navigation, and footer.