AI Ed Wiki logoAI Ed WikiUse with AI

Synthesis: Maier, Seibold, and Klotz (2026) conduct a qualitative systematic literature review following PRISMA 2020 guidelines, synthesizing 47 peer-reviewed empirical studies (2023–2026: 3 in 2023, 13 in 2024, 24 in 2025, 7 in 2026) containing 121 distinct research questions on LLM-generated formative feedback. The search covered ERIC, Scopus, IEEE Xplore, SpringerLink, and ScienceDirect (an initial November 2025 pass yielded 268 unique records; an expanded February 2026 pass produced 638 unique documents, reduced to 429 peer-reviewed journal articles, 69 full texts examined, and 47 retained). Using deductive and inductive qualitative content analysis (Mayring), the review maps three research questions spanning theory, methodology, and empirical findings. The literature is grounded in five broad theoretical families — formative Feedback theory, Self Regulated Learning and motivational–affective theories, general Learning Theories, domain-specific frameworks, and technology-acceptance/human–AI frameworks — and is dominated by short-term experiments, quasi-experimental classroom studies, and expert evaluations, mostly in Higher Ed. Most studies use proprietary GPT-3.5/GPT-4 with context-enriched, role-based zero-shot or few-shot prompting; fine-tuned, open-source, and multi-agent approaches are emerging. Empirically, LLM feedback consistently beats no-feedback conditions, improving revision quality, Motivation, and short-term learning, sometimes approaching teacher feedback under well-designed prompting — but recurring risks (hallucinations, over-positivity, misclassification) call for mitigation via instruction fine-tuning, grounding prompts in student artifacts, and teacher-in-the-loop oversight.

Key Findings

Scalable formative feedback. Since late 2022, instruction-tuned LLMs have enabled scalable generation of elaborated, contextualized formative feedback for text-based student assignments across domains, moving beyond correctness-based responses to higher-order, criterion-based guidance.

Search and selection. Five databases were searched; after deduplication 638 unique documents were reduced to 429 peer-reviewed journal articles, of which 90 abstracts flagged by Elicit Pro semi-automated screening were manually reviewed and 69 full texts examined; 22 were excluded (lacking empirical feedback data, non-educational domains, synthetic datasets), leaving 47 studies from groups mainly in the USA, China, and Germany.

Theoretical grounding. Five theoretical families anchor the field: formative feedback theory (Hattie & Timperley's "Where am I going? / How am I going? / Where to next?" used as an evaluation rubric), SRL and motivational–affective theory (Self Determination Theory, expectancy-value, control-value), general learning theories (Vygotsky's ZPD, Cognitive Load Theory), domain-specific frameworks (SLA and Written Corrective Feedback, Toulmin's argumentation), and technology-acceptance/human–AI frameworks.

Methodological profile. Studies are dominated by short-term experiments, quasi-experimental classroom studies, and expert evaluations of feedback quality, mostly in Higher Ed; feedback was examined along three dimensions — feedback quality, feedback processing, and Learning Gains.

Technological trends. Most studies deploy proprietary, pre-trained GPT-3.5/GPT-4 with context-enriched, role-based zero-shot or few-shot Prompt Engineering; instruction fine-tuning, open-source models, and multi-agent designs are beginning to emerge.

Empirical findings. LLM-generated feedback consistently outperforms no-feedback conditions and often improves revision quality, motivation, and short-term learning outcomes, sometimes approaching teacher feedback under well-designed prompting — central to the wiki's Feedback and AI Feedback Quality concepts.

Recurring risks and mitigations. Hallucinations, over-positivity, and misclassification of student work recur; emerging evidence indicates that instruction fine-tuning, grounding prompts in student artifacts, and teacher-in-the-loop oversight can mitigate these issues.

Connected Concepts

Connected Articles

Citation

Maier, U., Seibold, M., & Klotz, C. (2026). LLM-generated formative feedback in education: A qualitative systematic literature review. Computers and Education Open, 100374. https://doi.org/10.1016/j.caeo.2026.100374