🏷️ formative-assessment
89 pages tagged with formative-assessment(78 articles, 11 concepts)
📄 Making AI-Generated Feedback Matter: From Provision to Student Enactment
> **Synthesis:** Alsaiari et al. (2026) report a large-scale quasi-experimental cohort study (13,037 students; 51,296 student-authored resources) comparing three AI-mediated feedback workflows. Studen…
2026-08-13 · feedback-loop, learning-analytics, higher-ed, student-experience, self-regulated-learning
📄 Peer and AI Review + Reflection (PAIRR): A Human-Centered Approach to Formative Assessment
> **Synthesis:** Sperber et al. (2025) present the Peer and AI Review + Reflection (PAIRR) model, a human-centered approach to formative assessment that combines peer review best practices with AI rev…
🏷️ Peer Review
> **Peer review** — the practice in which students read, evaluate, and provide feedback on one another's work, most often writing. In writing pedagogy, peer review is a long-standing best practice: st…
📄 Reimagining feedback through generative AI in engineering education
> **Synthesis:** Pecuchova, Benko, and Drlik (2026) investigate the capacity of a large language model (GPTo1) to generate formative feedback for student-created UML diagrams in a university software …
2026-08-10 · generative-ai, ai-feedback-quality, higher-ed, curriculum-design, self-regulated-learning
🏷️ AI Feedback Quality
> **AI feedback quality** — the accuracy, usefulness, timeliness, and pedagogical value of feedback generated by AI systems for learners. As AI-generated feedback becomes ubiquitous in education, unde…
🏷️ Assessment Validity in AI Education
> **Assessment validity** — whether assessments measure what they claim to measure. AI in education raises fundamental validity questions: do AI-graded assessments assess student learning or AI prompt…
🏷️ Automated Assessment
> **Automated assessment** — the use of AI to evaluate student work, from formative quizzes to high-stakes exams. Automated assessment spans multiple modalities — multiple-choice, short answer, essay,…
🏷️ Automated Grading
> **Automated grading** — AI systems that evaluate student work, from multiple-choice scoring to essay assessment and code review. Automated grading is one of the most mature and widely-deployed AI in…
🏷️ Feedback Loop
> **Feedback loop** — the cyclical process where AI systems assess student work, deliver feedback, observe the student's response, and adapt subsequent instruction. Effective feedback loops close the …
2026-08-09 · ai-feedback-quality, automated-grading, scaffolding, self-regulated-learning, ai-tutoring
🏷️ Learning Analytics
> **Learning analytics** — the measurement, collection, analysis, and reporting of data about learners and their contexts for the purpose of understanding and optimizing learning. AI has transformed l…
📄 Agentic AI and Pedagogical Best Practice: The Tension Between Automation and Learning
> Education AI is shifting from passive chatbots to **proactive agents** that initiate and pursue goals. This offers personalisation but risks undermining **learner agency and cognitive effort**. The …
📄 From authentic products to authenticated processes: authentic assessment in AI-rich higher education
> Generative AI has not created the need for authentic assessment — it has made weaknesses in assessment design harder to ignore. Polished products can now be generated or substantially mediated by to…
📄 CoTAL: Human-in-the-Loop Prompt Engineering for Generalizable Formative Assessment Scoring and Feedback
> 1. **Evidence-Centered Design (ECD)** — assessments and rubrics aligned to curriculum goals from the start 2. **Human-in-the-loop prompt engineering** — labelled examples and prompts refined iterati…
📄 Responsible Assessment in the AI Era: Key Insights from a Future-Focused Conference
> **Responsible assessment in the AI era** — assessment grounded in learners' sociocultural contexts and designed to generate valid, trustworthy, context-specific inferences from accumulated evidence,…
📄 ICLE++: Modeling Fine-Grained Traits for Holistic Essay Scoring
Introduces ICLE++, a new annotated corpus of persuasive student essays that addresses critical limitations of the dominant ASAP benchmark in [[automated-essay-scoring]] research. Unlike ASAP — used by…
2026-07-31 · automated-essay-scoring, automated-grading, benchmark, educational-measurement, higher-ed
📄 Designing Needs- and Attention-Aware AI Learning Tools for Engineering Education: Insights from Psychological Outcomes
Survey of 206 engineering students: AI chatbots provide greatest perceived benefit as relief from competence frustration, smaller benefits for autonomy, weakest for relatedness. Baseline motivational …
2026-07-30 · higher-ed, stem-education, student-experience, affective-computing, personalized-learning
📄 The Easy Trap: Why LLMs Underestimate Misconception-Driven Difficulty
LLMs systematically underestimate the difficulty of misconception-driven items ('The Easy Trap'). While LLM ratings show moderate rank correlation with empirical student difficulty (rho=0.52-0.70), th…
📄 Aligning LLM-Simulated and Human Examinees for Psychometric Calibration: A Cognitive Diagnostic Profiling Approach
Proposes Cognitive Diagnostic Profiling (CDP), a zero-shot framework that dramatically improves LLM-simulated examinee alignment with human test-takers. With CDP, IRT difficulty Spearman correlations …
📄 Archetypes or ability? Clustering for modelling student mathematical competence
On 119,034 students across 13 UK national exams, Bernoulli Mixture Models found few distinct skill clusters — overall ability dominates. A simple explainable model achieved 78% accuracy, competitive w…
📄 AICoFE: AI-Powered Feedback System
> **AICoFE** (AI-based Collaborative Feedback) is a multi-LLM feedback generation system for higher education that combines independently fine-tuned language models with **teacher-in-the-loop mediatio…
📄 The Effect of High-Frequency, Automatically-marked Formative Assessments on Student Outcomes in A-Level Sciences
This quasi-experimental mixed-methods longitudinal study (N=142) deploys a fully automated marking pipeline for handwritten mock examinations in A-Level sciences, removing the human-marking bottleneck…
📄 What Does the Credential Still Certify? Cognitive Stewardship for AI-Mediated Education
Generative AI undermines a basic premise of educational assessment: that submitted work reliably evidences the human capacities a credential certifies. This paper proposes *cognitive stewardship*, a f…
📄 Assessment in Team Problem-Solving Exercises in Computing Education
Tabletop exercises (TTXs) let learner teams rehearse high-stakes workplace tasks such as cybersecurity incident response, but their open-ended, collaborative nature makes [[formative-assessment]] diff…
📄 Student Evaluation of Repeated AI Feedback Across a Semester of Writing
This short paper provides rare descriptive classroom evidence on what happens when students repeatedly use generative-AI feedback across a full semester of writing coursework. Drawing on 2,988 reflect…
📄 Artificial intelligence and feedback in university education: effectiveness and student perceptions
This quasi-experimental study directly compares **AI-generated feedback** (two LLMs: **GPT-o4-mini** and **DeepSeek R1**) with **expert human-teacher feedback** in a project-based university course (A…
📄 Benchmarking Multimodal Large Language Models for Scientific Visualization Literacy
Multimodal large language models (MLLMs) are increasingly used to interpret visualizations, yet most evaluations remain chart-centric and offer limited insight into **scientific visualization (SciVis)…
📄 Learning Engagement Assistant (LEA): Cross-Course Scalability and Classroom Evaluation of an Agentic AI Tutoring System
LEA (Learning Engagement Assistant) is an **agentic AI tutoring system** that couples course-specific retrieval-augmented generation (RAG) with structured [[knowledge-tracing]] / Knowledge Component (…
📄 LLM-Generated Design Problems for Assessing Higher-Order Thinking in Project-Based Learning
Introduces 'design problems' (DPs): concise, scenario-based prompts that require applying knowledge in transfer contexts, generated with LLMs to assess higher-order thinking (HOT) in project-based lea…
📄 A Durability and Cross-Language Transfer Benchmark for a Validated Teaching-Feedback Classification Protocol
> Extends a prior validated protocol for classifying open-ended teaching-evaluation feedback by thematic category and sentiment, introducing a durability and cross-language transfer benchmark. Institu…
📄 AIED's Unfinished Mission: Centering Agency and Motivation in the Age of Effortless Bypass
The widespread availability of general-purpose AI that can perform complex cognitive tasks threatens to undermine education at scale. This effortless bypass dilemma sharpens a challenge AIED has long …
2026-07-09 · over-reliance, student-experience, self-regulated-learning, metacognition, teacher-role
📄 DebugTracker: Lightweight Process Evidence for Classroom Debugging
Debugging exercises are usually graded from final code and test outcomes, which hide *how* students reproduced failures, formed hypotheses, inspected evidence, edited code, and verified fixes. The aut…
📄 Prompt Coach: An Empirical Evaluation of an Agentic Tutor for Learning Prompt Engineering in Software Development
Prompt engineering is a critical yet undertaught skill for software developers, poorly served by traditional instruction because of its evolving, interactive, context-dependent nature. The authors int…
📄 When AI Is Wrong on Purpose: How Students Respond to Buggy GenAI Code
As generative AI becomes central to software development, CS education is shifting toward prompt-centered workflows where students describe intended behavior in natural language to elicit code. But pr…
📄 Automated Grading of Linux/Bash Examinations Using Large Language Models
**Manuel Alonso-Carracedo, Ruben Fernandez-Boullon, Pedro Celard, Francisco J. Rodriguez-Martinez, Lorena Otero-Cerdeira (2026)** This paper presents an [[llm]]-based grading system for Linux/bash com…
📄 Constructing Epistemic AI Literacy: Detecting Epistemic Aims and Processes in Student-AI Co-Programming
**Mengqian Wu (2026)** Epistemic thinking — understanding how knowledge is constructed and justified — plays a central role in [[ai-literacy]], particularly when students co-program with generative AI…
📄 Data Comics for Education: Evaluating Effectiveness, Benefits, and the Ethics of AI-Assisted Creation
Data comics combine sequential visual narratives with data visualization to improve student engagement with [[generative-ai]] in educational settings. This paper evaluates the effectiveness of AI-assi…
📄 Evaluating Interactivity: Toward Automated Assessment of AI-Generated Explorable Explanations
While [[llm]]s now enable rapid generation of learning materials like [[generative-ai]], evaluating the pedagogical quality of these materials remains an open challenge. This paper proposes an automat…
📄 From Answer Generators to Reasoning Facilitators: Designing AI Tutors for Mathematical Reasoning in High-Stakes Environments
The rapid integration of [[llm]]s into [[intelligent-tutoring]] threatens to reduce mathematical learning to mere answer generation. This paper presents a design framework for AI tutors that act as re…
📄 CogTax: A Four-Level Cognitive Taxonomy for Command-Line Computing Education
> **Manuel Alonso-Carracedo, Ruben Fernandez-Boullon, Pedro Celard, Francisco J. Rodriguez-Martinez, Lorena Otero-Cerdeira** — Universidade de Vigo, submitted 30 Jun 2026…
📄 To Tab or Not to Tab: Measuring Critical Engagement in AI Code Completion Tools Using Behavioral Signals and Attention Checks
Hutchison et al. (2026) develop and validate a method for measuring critical engagement with AI code completion tools in educational settings. Using behavioral signals (time-to-accept, edit distance f…
📄 AI Coaching for Accelerating Human Skill Development with Reinforcement Learning
This paper explores how an embodied AI agent can act as a [[scaffolding|coach]] that accelerates human motor-skill development using [[adaptive-learning|reinforcement learning]]. The authors argue tha…
📄 Cross-Subject Predictive Validity for Learning Outcomes of Delayed Start Behavior
This study examines the [[student-modeling]] validity of **delayed start behavior** — when students begin assignments or practice sessions past a recommended start time — as a predictor of learning-ga…
📄 Do Gains from Generative AI-Enabled Adaptive Pretesting Persist? Evidence from a Retention Study
Akgun and Toker (2026) examine whether learning gains from GenAI-enabled adaptive pretesting persist over a seven-week retention period. Undergraduate participants completed adaptive AI-assisted prete…
📄 The Correct Answer Trap: Pedagogically-Grounded Detection and Feedback for Hidden Misconceptions
Imran and Bulathwela (2026) identify the 'correct answer trap' — automated feedback systems that judge only answer correctness reinforce rather than address misconceptions when students reach the righ…
📄 Why Machines Misread Pedagogical Quality: Human-Machine Alignment in LLM-Based Pretest Question Evaluation
Tseng et al. (2026) investigate human-machine alignment in LLM-based pretest question evaluation — a critical bottleneck for scalable AI-assisted assessment. Their AI-assisted workflow combines automa…
📄 Estimating Learners' Skill Acquisition Without Temporal Information
Nagai et al. (2026) tackle the practical problem that many real-world educational datasets contain only single-time-point assessments (snapshots) without temporal information, making standard time-ser…
📄 Confidence-Aware Automated Assessment of Student-Drawn Scientific Models
> **Luyang Fang, Yingchuan Zhang, Jongchan Park, Zhaoji Wang, Ping Ma, Xiaoming Zhai** (2026). arXiv cs.AI preprint…
📄 PsyScore: A Psychometrically-Aware Framework for Trait-Adaptive Essay Scoring and ZPD-Scaffolded Feedback
> **Wei Xia, Jin Wu, Haoran Shi, Xiangyu Wang, Chanjin Zheng** (2026). East China Normal University / arXiv cs.CL preprint…
📄 AI-Driven Assessment of Human Tutors: Linking Training Performance to Real-Life Practice
AI-driven assessment of human tutor training performance correlates with real-life tutoring quality; bridges the gap between training metrics and classroom practice. AI-Driven Assessment of Human Tuto…
📄 Cross-Dataset Bloom Question Classification: Supervised Models and Prompted LLMs
Evaluates cross-dataset generalization of ML/DL methods and LLMs for automatic Bloom's taxonomy classification of assessment questions across five datasets. Supervised ML/DL models degraded substantia…
📄 AI-Integrated Learning Management System for Middle School: A Longitudinal Study of Learning Outcomes
**Misan Paul Etchie, Taiwo Olutosin** — cs.CY, cs.AI, cs.HC This paper proposes an AI-integrated LMS designed specifically for middle school instruction, addressing the gap between current LMS platfor…
📄 Hybrid E-Assessment in Higher Education: Semi-Automated Grading of Paper-Based Written Examinations
**Hartwig Grabowski, Michael Canz** — cs.AI, cs.CV, cs.CY This paper identifies the didactic narrowing caused by fully digital e-assessment (overuse of closed question formats) and proposes a hybrid a…
📄 TibetCPR: A Multimodal Tactile Feedback System for CPR Training in High-Altitude Regions
**Yibo Meng, Ruiqi Chen, Zhiming Liu, Xiaolan Ding** — Accepted at MobileHCI 2026 — cs.HC TibetCPR is a low-cost, self-guided CPR training system that pairs depth-driven electrotactile feedback with r…
📄 FOXGLOVE: Comparing Goal-Oriented Writing Feedback from Experts and LLMs
Introduces **FOXGLOVE**, a dataset of 696 feedback comments by trained writing instructors on 69 twelfth-grade argumentative essays, paired with 1,644 comments from four frontier LLMs — totaling 2,340…
📄 LLM-Generated Feedback in Introductory Programming: A Classroom Study
Presents a **large-scale classroom study** (N=215 students, 6,693 submissions across 17 labs) deploying AI-generated feedback through a randomized protocol in an introductory Python programming course…
📄 VISMATIC: Secure Containerized Framework for Process-Oriented CS Education Monitoring
Addresses a critical tension in [[stem-education|CS education]]: the widespread adoption of generative AI makes it impossible to distinguish authentic student effort from AI code synthesis by evaluati…
📄 Teacher-Authored Prompts for Configuring Student-AI Dialogue: K-12 Classroom Implementation
This large-scale K-12 deployment provides empirical evidence that teacher-authored prompts can reliably shape the cognitive quality of student-AI dialogue at classroom scale. The TASD system lets teac…
📄 The Tutoring Effectiveness Index: Predicting LLM Math Tutor Quality from Four Conversation Signals
> **Authors:** Shim Jaechang, Unggi Lee (2026) — CIKM 2026…
🏷️ AI Ed Evaluation
> **AI-ed evaluation** — the body of methods, benchmarks, and criteria used to assess whether AI education tools (LLM-based tutors, automated graders, feedback systems, agents) actually work — not jus…
📄 Catching The Correct Answer Trap: Characterising AI Tutor Blind Spots When Analysing Student Reasoning
**Catching the Correct Answer Trap** — accepted at AIED 2026 — exposes a critical blind spot in [[intelligent-tutoring]] systems: they systematically fail to detect misconceptions when students arrive…
📄 Mathematical Modelling of Ethical AI Use in Higher Education: A Coordination Game Framework for Future-Facing Learning
**Ethical AI Use in Higher Education: A Coordination Game Framework** provides a formal mechanism-level account of why policy statements alone fail to change student AI-use behavior. Reframing student…
📄 LLM-assisted sentiment analysis for integrated computational and qualitative mixed methods education research: A case study of students' written reflection assignments
**LLM-Assisted Sentiment Analysis for Mixed-Methods Education Research** demonstrates how LLMs can serve as scalable qualitative research assistants, enabling researchers to investigate multiple demog…
📄 REC-CBM: Rubric-Aware Error-Correction Concept Bottleneck Models for Trustworthy Open-Ended Grading
**REC-CBM: Rubric-Aware Error-Correction Concept Bottleneck Models** advances the [[automated-grading]] frontier by solving a fundamental trust problem: even accurate AI graders are unusable if educat…
📄 Slide Deck Q&A Quality Assurance App: A Multi-Stage Pipeline for Pedagogical Question Generation
SlidesQAQA is a Flask-based system that extracts text and rendered images from PDF lecture slides and processes them through a four-stage [[llm]] pipeline: **window planning** (segment extraction), **…
📄 Automated Grading of Handwritten Mathematics Using Vision-Capable LLMs
Automated grading systems have enabled scalable assessment for many response types, but handwritten mathematics remains a barrier due to the complexity of multi-step solutions. Vision-capable large la…
📄 Combating Harms of Generative AI in CS1 with Code Review Interviews and a Flipped Classroom
Oral code reviews paired with a flipped classroom represent a pragmatic harm-reduction approach to generative AI in CS education. Rather than banning LLMs, Fowles et al. (2026) designed weekly formati…
📄 What Makes Words Hard? Sakura at BEA 2026 Shared Task on Vocabulary Difficulty Prediction
🔗 [Code](https://github.com/adno/vocabulary-difficulty) This paper presents two complementary approaches to predicting vocabulary difficulty for language learners, achieving state-of-the-art results …
📄 Generative AI Feedback, English Writing and Teacher Rubrics: A Multiple-Case Study of CyberScholar
RAG-based rubric-grounded GenAI writing feedback improved student revision quality (N=143, grades 7-11) and saved teacher time, but automated ratings were inconsistent. CyberScholar demonstrates rubri…
📄 What Don't You Understand? Using Large Language Models to Identify and Characterize Student Misconceptions About Challenging Topics
This paper presents a systematic two-stage methodology for surfacing student misconceptions at scale. Drawing on 3,802 medical student enrollments across 5 biomedical science courses (9 course periods…
📄 Retrieval-Augmented Tutoring for Algorithm Tracing and Problem-Solving in AI Education
KITE (Knowledge-Informed Tutoring Engine) introduces a [[intelligent-tutoring]] architecture that grounds its responses in course materials through a multimodal [[scaffolding|RAG pipeline]]. Unlike ge…
📄 LaTA: A Drop-in, FERPA-Compliant Local-LLM Autograder for Upper-Division STEM Coursework
> LaTA: A Drop-in, FERPA-Compliant Local-LLM Autograder for Upper-Division STEM Coursework **Rodríguez (2026)** — Oregon State University. Submitted to Computers & Education.…
📄 LLM-based Multimodal AI Feedback Produces Equivalent Learning and Better Student Perceptions than Educator Feedback
**AI multimodal feedback matches educator feedback for learning while significantly outperforming it on student perceptions.** The authors built a real-time AI-facilitated multimodal feedback system i…
📄 Assessing the Impact and Underlying Pathways of Sequenced AI Feedback on Student Learning
**Sequenced AI feedback harms learning despite boosting engagement and positive perceptions.** In a randomized experiment with 199 participants, the authors compared two types of AI-generated feedback…
📄 AISSA: AI-based Student Slides Analysis Tool for Academic Presentations
> A web-based system that uses LLMs and Learning Analytics dashboards to provide automated, rubric-based feedback on student presentation slides. Developed by Becerra et al. (2026), AISSA addresses th…
📄 AI-Generated Lesson Plans in Civic Education
> An analysis of 310 AI-generated lesson plans (2,230 individual activities) produced by ChatGPT (GPT-4o), Gemini (1.5 Flash), and Copilot (GPT-4 based) for all 53 Massachusetts eighth-grade civics st…
📄 Prober.ai: Gated Inquiry-Based Feedback via LLM-Constrained Personas for Argumentative Writing
> A web-based writing environment that inverts the AI-tutoring paradigm: rather than generating improved text for students, Prober.ai constrains an LLM to ask only targeted inquiry-based questions abo…
📄 Engagement Assessment in Video Learning
> **EduGage** (Leng et al., 2026) addresses a core challenge: in online/video-based learning, **learners must self-regulate** their engagement with instructional materials. > Sensor-based momentary as…
📄 Programming Intelligent Tutoring Systems
> **SCRIPT** (Deriyeva, Dannath, Paassen, 2026) implements an intelligent tutoring system for **Python programming** in a German university context, filling a gap in prior ITS which rarely supported P…
📄 Quality-Conditioned Agreement in Automated Short Answer Scoring: Mid-Range Degradation and the Impact of Task-Specific Adaptation
> Schleifer, Ariely & Klebanov (2026) investigate a critical gap in [[automated-grading]]: **how scoring quality degrades for mid-range student responses**. Most ASAS evaluations focus on clearly corr…
📄 TeachBench - Evaluating LLM Teaching Ability
> While LLMs are increasingly used as teaching assistants, their teaching capability remains insufficiently evaluated — a critical gap in current AIED research. > Syllabus-grounded framework for measu…
📄 AI Peer Feedback Systems
> Peer feedback develops critical reflection and evaluative judgment, yet: > Student peer feedback is often superficial or inconsistent. **AICoFe** (AI-based Collaborative Feedback) uses a multi-LLM p…
📄 Authentic Assessment
> Wiggins (1990) proposed AA as a counterbalance to standardised tests: direct examination of "student performance on worthy intellectual tasks." > Authentic assessment (AA) has evolved from workplace…
📄 Automatic Short Answer Grading with LLMs
> Automatic Short Answer Grading (ASAG) is never perfect. Upper bounds on accuracy arise from: > Zero-shot LLMs perform strongly on ASAG without task-specific fine-tuning, but **model-based confidence…
📄 Collaborative AI Tutoring
> ProPACT constructs a real-time model of pair collaboration using three signals: > Most adaptive learning systems are individual-centric and reactive. **ProPACT** treats **collaboration itself as the…
📄 Educational VLM Evaluation
> Benchmarking vision-language models (VLMs) not on their ability to solve problems, but on their ability to *support learners* — particularly struggling learners and those making errors. Traditional …
📄 From Surface Learning to Deep Understanding: A Grounded AI Tutoring System for Moodle
> Ostrowska, Kukla & Majstrak (2026) present an AI tutoring system **integrated into the Moodle LMS** designed to scaffold students from surface-level fact recall to deep conceptual understanding thro…
🏷️ Metacognition
> Metacognition — thinking about one's own thinking — is both a target of AI education research (can AI tools develop students' metacognitive skills?) and a risk factor (AI completing tasks may suppre…
🏷️ Self-Regulated Learning
> Self-regulated learning (SRL) describes learners as active participants who can shape and develop their cognitive and behavioral actions in a successful way. AI tools can either scaffold SRL develop…
🏷️ Socratic AI Dialogue
> Socratic dialogue — asking structured questions rather than providing answers — is one of the strongest pedagogical scaffolds for deep learning. When automated via AI, it produces measurable reasoni…