Concept
Formative Assessment
Formative assessment — assessment designed to inform ongoing instruction and learning, as opposed to summative evaluation. In AI education, formative assessment is both transformed by AI and essential to it: AI systems can generate, validate, and adapt formative items and feedback at scale, while formative feedback is a primary mechanism through which AI tutors and adaptive systems support learning. The knowledge base's research examines AI-generated formative items, AI-generated feedback, and the design and evaluation of these systems.
Questions to Consider
- Formative assessment is meant to close the loop — surface what students don't know and give feedback they can act on while learning is still in progress. How is that fundamentally different from summative evaluation, and when might the two get confused in practice?
- AI can generate multiple-choice questions with impressive accuracy on verifiable dimensions, but is weakest on instructional-judgment dimensions. If a machine is good at correct answers but weaker at pedagogical judgment, what should it be trusted to do — and what should humans keep doing?
- AI-generated feedback only helps when students actually enact it — the 'enacted feedback' condition, where students select, evaluate, and apply suggestions, outperformed simply being given feedback. If enactment matters more than the feedback itself, what does that mean for how formative feedback should be designed?
- Feedback is described not as information transfer but as an ethical, relational practice. What gets lost when formative assessment is mass-produced as 'AI slop' — and what can human comment banks and relational care preserve?
Introduction
Formative assessment is central to AI in education because it sits at the junction of Assessment and learning feedback. Its purpose is to close the loop: surface what students know and don't know, and provide feedback they can act on to improve. AI makes this feasible at scale — generating items, scoring responses, and delivering individualized feedback — but the knowledge base's research shows that quality varies dramatically across item types, and that feedback only helps when students actually enact it.
AI-generated formative items
AI systems generate formative assessment items across modalities, with reliability varying by type:
- Multiple-choice questions: CODE-GEN: A Human-in-the-Loop RAG-Based Agentic AI System for Multiple-Choice Question Generation shows agentic AI can reliably generate MCQs for coding comprehension when validated across seven pedagogical dimensions — success rates reach 98.6% for concept alignment and 79.9% for feedback quality — suggesting AI is strongest on verifiable dimensions and weakest on instructional-judgment dimensions. This connects to automated question generation more broadly.
- Automated essay scoring: multi-agent frameworks (e.g., MASS) improve consistency over stand-alone LLMs for essay scoring, though interpretability of multi-agent scoring decisions remains an open challenge.
- Formative scoring pipelines: CoTAL couples Chain-of-Thought prompting with active learning and Evidence-Centered Design to produce generalizable formative-assessment scoring with human-in-the-loop prompt engineering.
- High-frequency, automatically-marked assessments: automated formative assessments in A-level sciences examines the effect of high-frequency, automatically-marked formative assessment on learning outcomes. A scoping review of short-answer auto-marking in science (2017–early 2024) confirms this formative short-answer use case is a field with real traction: BERT-family models dominated auto-marking through 2021 before prompting larger LLMs from ~2022, and domain-augmented, rubric-aware, and chain-of-thought models performed best — yet the review's calls for comprehensive evaluation and unresolved fairness and explainability gaps caution against extending such systems to unmediated summative or high-stakes use (Auto-marking short answer questions in science: The foundational years of transformer-based models from BERT to GPT-4).
AI-generated feedback
A large body of knowledge base research examines AI-generated formative feedback:
- The enactment problem: Making AI-Generated Feedback Matter (13,037 students; 51,296 resources) shows feedback value depends on whether students enact it — the Enacted Feedback condition, where students select, evaluate, and apply AI feedback suggestions, outperformed simple directed feedback.
- Feedback is not information transfer: The care-full craft of feedback argues feedback is an ethical, relational practice, not information transmission — feedback only constitutes feedback when students make sense of and act on it, and contrasts mass-produced "AI slop" with human comment-bank shortcuts.
- Sequenced feedback can backfire: Sequenced AI feedback (encouragement → hints → correct answer, designed to promote autonomy) actually harmed learning in a randomized experiment (N=199) despite boosting engagement and positive perceptions — a cautionary finding about feedback design.
- Learner-centered tools: PolyFeed combines ML suggestion models with teacher practice, showing how teachers adopt and adapt AI feedback suggestions; AI-supported internal feedback helps undergraduates develop evaluative judgment.
- Rubric-guided prompting and role-aware feedback: Yaşar et al. (2026) showed that iterative rubric co-refinement — clarifying performance descriptors and explicitly accepting implicit indicators of learning — drove LLM–human agreement on student design work from 54.75% to 81.25% (Cronbach's Alpha rising from 0.393 to 0.798), with the largest gains in the cognitively demanding Iteration & Reflection category. Prompting the same model under instructor, peer-reviewer, and grant-reviewer roles produced distinct evaluative feedback, and post-revision LLMs were more consistent than some human raters in applying performance thresholds — positioning rubric-guided LLMs as calibration and co-design partners in formative feedback environments, with human-in-the-loop oversight remaining essential.
- AI feedback sustains participation and drives gains at scale: Geschwind et al. (2026)'s semester-long field experiment across undergraduate tutorials found that individual GPT-4 formative feedback (spanning all three Hattie & Timperley dimensions — Feed-Back, Feed-Up, Feed-Forward) sustained the highest participation across eight open-ended tasks, lengthened student answers, and produced the strongest content learning gains — an effect driven by reliable, consistent AI provision, since when high-quality textual peer feedback was actually received, peer outcomes matched AI's.
- Feedback futures: Feedback Futures synthesizes a special issue and argues the question is not whether GenAI can produce feedback but how to design feedback that supports learning, distilling recurring tensions across the field.
- Diagnosis-first feedback for open-ended quantitative problems: Arthur (Yin et al. 2026) delivers real-time, personalized formative feedback on Engineering Economics Calculated Formula Questions, a domain where handwritten, unstructured solutions had previously blocked AI support. A per-question XGBoost backbone diagnoses likely rubric-labeled mistakes from students' submitted numerical answers (average precision 0.81, recall 0.79), and a dialogue-based scheme requests intermediate answers only when prediction confidence is low — balancing feedback accuracy against collection efficiency within a question-bank web interface.
- Scale and limits of LLM formative feedback (systematic evidence): a PRISMA-guided systematic review of 42 empirical studies (2023–2025) finds LLMs can reduce teacher workload and deliver rapid, personalized feedback at scale — especially in large or higher-education cohorts — but that feedback is sometimes too generic or misaligned with the assigned grade and reliability slips on longer, multilingual, or nuanced tasks, reinforcing that formative AI feedback is best deployed under educator oversight (Can ChatGPT Replace the Teacher in Assessment? A Review of Research on the Use of Large Language Models in Grading and Providing Feedback).
- Adaptivity is a separable ingredient, not decoration: Sukjaitham, Schaaf, Brod & Breitwieser (2026) supply the direct causal test that most LLM-feedback studies assume away, pitting GPT-4 response-contingent feedback against expert-written generic guidance matched on structure, tone, length, and motivational phrasing (verified with a five-dimension quality rubric, κ = .76–1.00). In a preregistered within-subjects experiment, 155 German fifth- and sixth-graders (M = 12.08 years) revised six if-then plans: plan quality rose from a median of 2 → 5 under adaptive feedback versus 2 → 3 under generic guidance (within-person V = 10,440, p < .001, r = .86; condition × time interaction estimate = 1.68, SE = 0.14, p < .001, with no pre-support difference). Children rated adaptive feedback both more helpful (r = .67) and more motivating (r = .74), and trial-level perceptions predicted the size of revision gains — making perceived usefulness part of the pathway rather than an affective byproduct. Because the control was itself well designed, the study shows added value beyond good non-contingent guidance rather than the difference between feedback and nothing: generic guidance is a genuine but limited substitute that plateaus at its own median. Planning served as the test case as a core self-regulated learning strategy with explicit quality criteria, which makes contingency — the property that distinguishes Scaffolding from static support — directly measurable on a one-sentence response. The authors define adaptivity narrowly as response-contingent adaptation of feedback content to the learner's concrete response, distinguishing it from conversational interactivity, tone, and stable-trait adaptive learning.
- Perceived usefulness tracks actionability: Mendonça et al. (2026) had 144 programming students rate 893 LLM-generated feedback instances on five dimensions and 237 consolidated reports on six, holding domain, task, and instrument constant while educational level varied. Ratings were favorable throughout, with student-level means of 4.24 to 4.43 for individual answers and 4.11 to 4.38 for reports, yet actionability and usefulness drew the lowest ratings even as actionability and perceived accuracy carried the largest relative weights (32.6% and 29.0%) in a model explaining 76% of the variance in perceived usefulness, and motivation and personalization led a report-level model of intention to use explaining 53%. Since actionability is the dimension Feedback research treats as hardest to provide, the pattern reads as technology acceptance applied to feedback: usefulness tracks whether a learner can act, not how polished the feedback sounds.
- Comprehension outruns actionability on the path to revision: Yao and Fan (2026) separated issue-level diagnosis from suggestions and gave students a written interpretation sheet (restatement, effect explanation, uncertainty identification, revision plan) before independent revision, structured teacher review, reassessment and reflection. Across three writing cycles in two intact classes (96 students; 288 cycle-level observations), feedback comprehension carried the strongest association with revision quality (r = 0.501, ahead of actionability at 0.480 and diagnostic accuracy at 0.445), and the diagnostic group gained 5.92 writing points against 3.58 (Hedges' g = 1.12), locating the leverage in students' understanding of the diagnosis rather than in its actionability alone.
Curriculum-grounded and educator-in-the-loop design
LearnLens addresses three persistent problems in AI formative assessment: error-aware assessment (capturing nuanced reasoning errors rather than surface mistakes), topic-linked memory chains (replacing noisy similarity-based RAG (Retrieval-Augmented Generation) with structured curriculum-grounded retrieval), and educator-in-the-loop design (teacher customization and oversight, not full automation). This connects to the broader tension in Human-in-the-Loop: scalable automation with expert validation.
Hoppe, Loibl & Leuders (2026) sharpen what educator-in-the-loop means once a tool produces its own claims rather than raw observations. Their conceptual analysis argues that AI-generated diagnostic inferences are qualitatively different evidence, because they are already the product of algorithmic interpretation, so teachers need a further layer the authors call meta-diagnosis: deliberately accepting, rejecting, or modifying an inference and integrating it with their own contextual knowledge. Specifying the DiaCoM framework for this treats AI-generated inferences as a situation characteristic and the accept, reject, or modify decision as diagnostic behavior, while expanding the person characteristics teachers need to include knowledge of how AI systems actually work. Because current systems rest mostly on performance data such as task correctness and completion time, motivational states and classroom dynamics stay largely absent, so the paper keeps teachers, not the dashboard, as the responsible reflective agents and frames judging algorithmic claims as a target for professional development.
Design trade-offs
| Dimension | AI Suitability | Human Requirement |
|---|---|---|
| Factual correctness | High | Low |
| Concept alignment | High | Medium |
| Distractor quality | Low | High |
| Feedback depth | Low | High |
| Rubric consistency | Medium | Medium |
Assessment, feedback, and learning
Formative assessment in AI education connects to the learning process itself:
- Feedback loops: feedback loops are the mechanism by which formative assessment informs learning; AI tutors and adaptive systems close these loops at scale.
- Self-regulated learning: formative feedback supports self-regulated learning when students monitor progress and adjust; AI feedback should cultivate evaluative judgment, not displace it.
- Automated scoring reaches some phases of the cycle, not all: Chen & Liu (2026) ran a 14-week quasi-experiment in which 46 interpreting students submitted weekly renditions to an automated scoring system that returned an immediate score, transcript, marked errors, and a reference rendition, while a control group received whole-class teacher feedback only. The automated group improved more overall (d = 1.03), but the gain stayed where the deficit was decomposable, the signal reliable, and the scale sensitive: linguistic accuracy and logical coherence rose while information fidelity (agreement with human raters r = 0.12) and delivery fluency did not move. Self-regulated learning was uneven in a matching way, with execution and monitoring correlating with score gains (r = 0.42) while planning and emotional motivation sat near the scale midpoint, and several students deferred engagement after low scores instead of analyzing causes. The system reached the performance phase of the cycle, not the planning that precedes it.
- From automated diagnosis to generated practice: Zhu, Luo & Li (2026) wire audio-score alignment, error quantification, and a Proximal Policy Optimization layer into one closed loop for music education, turning rhythm error signals (91.2% recall, 98.4% specificity) into rewards that steer generated practice tracks, and report a significant Group x Time interaction favoring the system group (beta = 0.52) across a 12 week quasi-experiment with 120 undergraduate music majors. The authors frame this as technical feasibility, and the practical reading is triage rather than assessment: precision of 89.7% means roughly one flagged rhythm error in ten is a false alarm, and the study's expert reviewers rated support for musical expression as the weakest area, so automated loops suit technical drills while expressive judgment stays with teachers.
- Scaffolding: Scaffolding and formative assessment work together — AI can provide just-in-time hints and prompts, though sequenced feedback research cautions against over-structuring.
- Validity and quality: the quality and validity of AI-generated formative items and feedback must be evaluated; AI Ed Evaluation provides the methods.
Risk: Assessment as surveillance
Formative assessment systems can shift from learning-support tools to behavior-monitoring infrastructure. The same data streams that enable adaptive tutoring can enable punitive tracking if AI Governance is weak. This connects to Privacy and student well-being, and argues for formative systems that support learning rather than surveil it.
Implications for AI in education
- Match item type to AI reliability: use AI for verifiable dimensions (concept alignment, correctness) and retain human judgment for instructional dimensions (distractor quality, feedback depth).
- Design for enactment, not just provision: AI feedback only helps when students select, evaluate, and apply it — structure workflows that support enactment.
- Feedback design matters more than volume: sequenced or over-structured feedback can backfire; prioritize feedback that supports student sense-making and autonomy.
- Keep educators in the loop: curriculum-grounded, educator-in-the-loop systems improve relevance and reduce noise.
- Evaluate quality and validity: assess AI-generated items and feedback for quality, validity, and equity, not just generation speed.
- Treat formative assessment as a philosophy, not a toolkit. Mesny, Roberge-Maltais & Galy (2026) synthesize the wider higher-education literature into an "assessment for learning" paradigm that frames formative, ongoing, and individualized Feedback as an overarching philosophy rather than a mere toolkit — balancing formative with summative purposes and foregrounding student agency, self-regulation, and metacognitive skill. They identify five mutually reinforcing practices (authentic assessment, self- and Peer Assessment, reassessment, standards-based grading, ungrading) through which this philosophy can be enacted, while noting their uptake remains highly uneven across higher-education fields.
Connected Concepts
- Assessment
- Educational Measurement
- Automated Assessment
- Automated Question Generation
- Assessment Validity
- Feedback
- AI Feedback Quality
- Feedback Literacy
- Self-Assessment
- Learning Analytics
- Personalized Learning
- Adaptive Learning
- Scaffolding
- Self-Regulated Learning
- Human-in-the-Loop
- Intelligent Tutoring
- AI Ed Evaluation
- Summative Assessment — Summative assessment: AI-resistant formats (oral, proctored, closed-book exams)
Connected Articles
- SSAIL: A Design Framework for Safe and Sound AI for Learning — SSAIL: A Design Framework for Safe and Sound AI for Learning
- Causal Modelling of Support Interventions for Student Competency Assessment — Causal Modeling of Support Interventions for Student Competency Assessment
- It Takes a Village... Program-Wide Approaches to Redesigning Assessment in a Time of Generative Artificial Intelligence (GenAI) — Program-wide approaches to redesigning assessment in the GenAI era
- Students' Perceptions of Multiliteracies Development Using AI-Assisted Portfolio Assessment — Students' perceptions of multiliteracies development with AI-assisted portfolio assessment
- Making AI-Generated Feedback Matter: From Provision to Student Enactment — Making AI-generated feedback matter: from provision to enactment
- The care-full craft of feedback in an age of generative AI — The care-full craft of feedback in an age of GenAI
- Feedback futures: beyond the limits of human and GenAI capacities — Feedback futures: beyond the limits of human and GenAI capacities
- LLMs Do Not Grade Essays Like Humans — LLMs do not grade essays like humans (Mathew et al. 2026)
- Assessing the Impact and Underlying Pathways of Sequenced AI Feedback on Student Learning — Impact and pathways of sequenced AI feedback
- Enhancing learner-centered feedback with AI: teachers'' practices and perceptions — Enhancing learner-centered feedback with AI
- From automated scoring to learning diagnosis: a mechanism study of AI-supported formative assessment in English writing — From automated scoring to learning diagnosis: a mechanism study of AI-supported formative assessment in English writing
- Unravelling undergraduates' development of evaluative judgments through AI-supported internal feedback — Developing evaluative judgments through AI-supported internal feedback
- CoTAL: Human-in-the-Loop Prompt Engineering for Generalizable Formative Assessment Scoring and Feedback — CoTAL: formative assessment scoring with human-in-the-loop prompting
- The Effect of High-Frequency, Automatically-marked Formative Assessments on Student Outcomes in A-Level Sciences — High-frequency automated formative assessment
- Artificial intelligence and feedback in university education: effectiveness and student perceptions — AI-generated feedback in higher education
- Designing Needs- and Attention-Aware AI Learning Tools for Engineering Education: Insights from Psychological Outcomes — LearnLens: curriculum-grounded AI feedback
- Comparing Generative AI and teacher feedback: student perceptions of usefulness and trustworthiness — GenAI vs. teacher feedback comparison
- Students' engagement with ChatGPT feedback: implications for student feedback literacy in the context of generative artificial intelligence — ChatGPT feedback and engagement
- AICoFe: Implementation and Deployment of an AI-Based Collaborative Feedback System for Higher Education — AI-coffee feedback framework
- CODE-GEN: A Human-in-the-Loop RAG-Based Agentic AI System for Multiple-Choice Question Generation — CODE-GEN: validated MCQ generation
- Responsible Assessment in the AI Era: Key Insights from a Future-Focused Conference — Responsible assessment in the AI era
- Designing for Authentic Assessment: A Scoping Review — Designing for authentic assessment
- Automated Grading of Linux/Bash Examinations Using Large Language Models — Automated grading of Linux/bash exams
- Instructor and AI Roles in the Chemistry Classroom: Future Science Teachers' Perceptions in a ChatGPT-Enhanced Formative Assessment — Instructor and AI roles in ChatGPT-enhanced formative assessment
- Reconsidering the Use of Oral Exams and Assessments: An Old Way to Move Into a New Future — Reconsidering oral exams as authentic, AI-resistant assessment
- Assessment twins: An approach for strengthening assessment validity in the age of generative AI — Assessment twins for strengthening assessment validity in the age of GenAI (Roe, Perkins & Giray 2026)
- A Hybrid Reasoning Framework for Artificial Intelligence Assessment Rubric Generation in Human and Automated Contexts: Evidence from an Undergraduate Programming Course — HARMOGEN-R: AI assessment rubric generation
- AI-assisted, instructor-supervised grading and feedback in higher education: Design and evaluation of an end-to-end pipeline — AI-assisted instructor-supervised grading and feedback
- Adaptive Scaffolding for Cognitive Engagement in an Intelligent Tutoring System — Adaptive ICAP scaffolding in an ITS (BKT vs DRL)
- Assessing students' DRIVE: A framework to evaluate learning through interactions with generative AI — DRIVE: assessing learning through GenAI interaction (DRI + Visible Expertise)
- ChatGPT Solves All Tested Qiskit Homework Assignments — ChatGPT solves Qiskit homework; autogradable design
- Beyond the Traceback: Using LLMs for Adaptive Explanations of Programming Errors — LLM adaptive explanations of programming errors
- From evaluation to emulation: LLMs as agents of iterative pedagogical design — LLMs as agents of iterative pedagogical design
- Auto-marking short answer questions in science: The foundational years of transformer-based models from BERT to GPT-4
- GPT-4 feedback increases student activation and learning outcomes in higher education
- Arthur: An artificial intelligence powered teaching assistant system for Engineering Economics class
- Innovative assessment and grading practices in higher education: A critical exploration for management educators
- Can ChatGPT Replace the Teacher in Assessment? A Review of Research on the Use of Large Language Models in Grading and Providing Feedback
- Adaptivity Makes Feedback Effective: Evidence From AI-Generated Feedback on Children's Plans — Adaptivity makes feedback effective: evidence from AI-generated feedback on children's plans
- Comparative analysis of peer group and AI-generated feedback in peer assessment: Insights into feedback quality and student perceptions in higher education — Comparative analysis of peer group and AI-generated feedback in peer assessment: Insights into feedback quality and student perceptions in higher education
- Impact of automated scoring on interpreting performance and self-regulated learning: evidence from a pedagogical experiment — Automated scoring, interpreting performance, and self-regulated learning (Chen & Liu 2026)
- Rethinking teachers’ diagnostic skills in AI-supported formative assessment: from diagnosis to meta-diagnosis — Teachers' diagnostic skills in AI-supported formative assessment: from diagnosis to meta-diagnosis
- Perceived usefulness and intention to use large language model-generated feedback across three educational levels: a user-centred study in programming — Perceived usefulness and intention to use LLM-generated feedback in programming across three educational levels
- Adaptive teaching assistance model combining generative AI and big data analytics — Adaptive teaching assistance combining generative AI and big data analytics in music education
- Student Use of LLMs and the Limits of AI-Generated Question Difficulty in Data Science Courses — Student Use of LLMs and the Limits of AI-Generated Question Difficulty in Data Science Courses
- Instructional Governance by Design: A Framework for AI in Computing Education — Instructional Governance by Design: A Framework for AI in Computing Education