🏷️ llm
350 pages tagged with llm(308 articles, 42 concepts)
📄 Making AI-Generated Feedback Matter: From Provision to Student Enactment
> **Synthesis:** Alsaiari et al. (2026) report a large-scale quasi-experimental cohort study (13,037 students; 51,296 student-authored resources) comparing three AI-mediated feedback workflows. Studen…
📄 Methodologies for Improving the Quality of AI Tutoring in K-12 Education
> **Synthesis:** Udeshi et al. (2026), the team behind **Khanmigo** (Khan Academy's K-12 AI tutor, launched 2023), describe the metrics they use to measure AI tutoring quality and student engagement, …
📄 CyberAGENTS: Structured Autonomy for Agentic Gamified Learning in Cybersecurity
> **Synthesis:** Hornung et al. (2026) present **CyberAGENTS**, an agentic framework for gamified cybersecurity learning that enables *structured autonomy* through ontology-guided validation, schema-g…
2026-08-13 · agentic-ai, computing-education, cs-education, pedagogical-safety, professional-training
📄 EduSim-LLM: An Educational Platform Integrating Large Language Models and Robotic Simulation for Beginners
> **Synthesis:** Lu and Zhang (2026) present EduSim-LLM, an educational platform that integrates large language models with robot simulation to make robotic control accessible to beginners. Recognizin…
📄 ELBench: A Multi-Dimensional Benchmark for Education-Facing Large Language Models
> **Synthesis:** Jiang et al. (2026) introduce **ELBench**, the first benchmark to evaluate education-facing LLMs on all four required dimensions — General Capability, Safety and Trustworthiness, Basi…
📄 Integrating Generative AI into Cybersecurity Education: A Study of OCR and Multimodal LLM-Assisted Instruction
> **Synthesis:** Patel et al. (2025) present an LLM-assisted instructional integration with a virtual cybersecurity lab platform, addressing workforce reskilling needs driven by the digital transforma…
2026-08-13 · generative-ai, cybersecurity, higher-ed, experiential-learning, instructional-assistant
📄 Would You Let a Humanoid Play Storytelling With Your Child? A Usability Study on LLM-Powered Narrative Human-Robot Interaction
> **Synthesis:** Lombardi et al. (2025) present a framework for enhancing the attention and social capability of the iCub humanoid robot by integrating advanced perceptual abilities that recognize soc…
📄 INSIDE the Student's Mind: Jointly Modeling Latent Reasoning and Action in LLM Student Simulators
> **Synthesis:** Niousha, Kang, & Norouzi (2026) introduce **INTERNAL STUDENT DIALOGUE (INSIDE)**, a student modeling framework that fine-tunes LLMs to both *act* like students and *think* like them. …
📄 ResidencyRL: Reinforcement Learning in Simulated Clinical Environments
> **Synthesis:** Liévin et al. (2026) present **ResidencyRL**, a reinforcement learning method for training clinical AI agents through simulated multi-turn clinical encounters (up to 60 dialogue turns…
📄 RoboBlockly Studio: Conversational Block Programming With Embodied Robot Feedback for Computational Thinking
> **Synthesis:** Li, Du, Sun, and colleagues (2026) design and evaluate RoboBlockly Studio, an integrated interactive system that combines block-based programming, a conversational AI teaching agent, …
2026-08-13 · computational-thinking, block-programming, educational-robotics, programming-education, k-12
📄 RoboBuddy in the Classroom: Exploring LLM-Powered Social Robots for Storytelling in Learning and Integration Activities
> **Synthesis:** Tozadore, Ertug, Chaker, and Abderrahim (2025) present RoboBuddy, an intuitive interface that lets teachers create scenario-based storytelling activities from their regular curriculum…
📄 Teachy Mini: Development and Preliminary Evaluation of a Knowledge-Based Generative Social Robot for Higher Education
> **Synthesis:** Vonschallen, Kaufmann, Oberle, Eyssel, and Schmiedel (2026) operationalize knowledge-based design (KBD) requirements for generative social robots (GSRs) by implementing them in the Re…
📄 AgentSchool: An LLM-Powered Multi-Agent Simulation for Education
> Ye et al. (2026) introduce **AgentSchool**, an LLM-driven multi-agent [[simulating-students|simulator]] that models learning as **state transition rather than prompted behavior**. It couples cogniti…
📄 ChatGPT-generated help produces learning gains equivalent to human tutor-authored help on mathematics skills
> Pardos & Bhandari (2024) report a randomized efficacy study (N=274) comparing ChatGPT-generated hints to human tutor-authored hints and a no-help control across four mathematics subject areas. Only …
📄 Multimodal Item Parameter Estimation using Simulated Response Probabilities
> **Synthesis:** This paper fine-tunes a multimodal large language model (Qwen3.5-based) to reconstruct multiple-choice model (MCM) and three-parameter logistic (3PL) item characteristic curves. By le…
📄 Embracing Imperfection: Simulating Students with Diverse Cognitive Levels Using LLM-based Agents
> Wu et al. (2025, ACL) tackle the core challenge of [[simulating-students]]: LLMs trained as "helpful assistants" produce overly perfect answers and fail to model the natural imperfections and varied…
2026-08-12 · simulating-students, generative-ai, student-modeling, knowledge-graph, cognitive-diagnosis
📄 Simulating Students with Large Language Models: A Review of Architecture, Mechanisms, and Role Modelling in Education with Generative AI
> Marquez-Carpintero, Lopez-Sellers & Cazorla (2025) present a thematic review of empirical and methodological studies using LLMs to [[simulating-students|simulate student behavior]] in education. The…
📄 Towards Valid Student Simulation with Large Language Models
> Yuan et al. (2026) present a conceptual and methodological framework for valid LLM-based [[simulating-students|student simulation]]. They identify the **competence paradox** — broadly capable LLMs a…
🏷️ Simulating Students
> **Simulating students** — using LLM-based agents to model learner behavior, cognition, and social dynamics for educational research, design, and training. Simulated students let researchers evaluate…
📄 Exploring AI-Supported Disciplinary Mediation in Student Project Teams' Text-Based Communication
> **Synthesis:** Cheng, Chung, Chiu, Lin & Liao (2026) present Spritz, a Discord-based [[llm]] technology probe that mediates disciplinary boundaries in interdisciplinary student project teams, findin…
📄 VeriForge: Mitigating Latent Knowledge Gaps in Narrative Drafting via Mixed-Initiative Scaffolding
> **Synthesis:** Sun et al. (2026) present VeriForge, a mixed-initiative [[generative-ai]] writing system that assumes initiative over domain discovery while the author retains initiative over narrati…
📄 A Bottom-Up Taxonomy of Student Discourse with a Socratic AI Physics Tutor
> **Synthesis:** Large language model (LLM) tutors are being deployed in introductory physics courses at a scale that produces transcript corpora far larger than traditional qualitative coding can abs…
📄 Adversarial Stress Testing of Role-Playing Language Agents using Multi-Agent Evaluation
> **Synthesis:** This paper presents a modular multi-agent platform for adversarially stress-testing [[agentic-ai|role-playing language agents]] through structured multi-turn dialogue. With three coor…
📄 Chat Debugging: An Exploratory Study of Human-AI Collaboration to Debug Analog Circuits
> **Synthesis:** This exploratory study investigates how undergraduates use [[llm|LLMs]] to debug malfunctioning analog circuits under exam conditions, identifying both promising [[human-ai-collaborat…
📄 Anchor Is the Key: Toward Accessible Automated Essay Scoring with Large Language Models Through Prompting
> **Synthesis:** Choi, Tate, Ritchie, Nixon & Warschauer (2025) investigate the most practical approach to LLM-based automated essay scoring — prompting — and find that providing anchor papers (exampl…
📄 Once a Response, Always a Response: Detecting LLM-generated Text via Latent Prompt Restoration
> **Synthesis:** EchoPrompt introduces a training-free zero-shot detector for [[plagiarism-detection|LLM-generated text]] that exploits the latent prompt dependency inherent in machine-generated conte…
📄 Does generative AI narrow education-based productivity gaps? Evidence from a randomized experiment
> **Synthesis:** In a [[rct|randomized controlled trial]] with 1,174 participants, Cruces et al. find that [[generative-ai|generative AI]] substantially narrows education-based productivity gaps, clos…
📄 Cognitive Offloading in Student–AI Collaboration: A Longitudinal Analysis of Prompting Strategies
> **Synthesis:** Misiejuk, López-Pernas, Kaliisa, and Saqr (2026) analyze 281 prompts from 122 student submissions across four assignments to examine how prompting strategies reveal cognitive offloadi…
2026-08-09 · cognitive-offloading, prompting-literacy, higher-ed, student-experience, learning-analytics
📄 ProPRL: Property-Aware Prerequisite Relation Learning in Educational Knowledge Graphs
> **Synthesis:** ProPRL advances [[adaptive-learning|prerequisite relation learning]] by going beyond conventional link prediction to adaptively integrate complementary educational evidence from conce…
2026-08-09 · adaptive-learning, knowledge-tracing, student-modeling, ai-education, personalized-learning
📄 Navigating the skill diversity frontier: How skill complexity explains worker resilience
> **Synthesis:** Using LinkedIn data on 2.4 million U.S. workers and 16,753 distinct skills, this paper introduces three complementary measures of skill complexity — specialization, diversity, and the…
📄 TACT: Taxonomy-Aligned Post-Training for Pedagogically Adaptive English Tutoring
> **Synthesis:** TACT (Taxonomy-Aligned Conversational Tutor) presents a human-grounded framework for training and evaluating pedagogically adaptive ESL tutors powered by [[llm|LLMs]]. Built on a Tuto…
📄 HiLLM-CD: LLM-Enhanced Hierarchical Cognitive Diagnosis
> **Synthesis:** Xie, Yang, Zhang, Li, Wang, Yang & Gao (2026) propose HiLLM-CD, a tree-structured framework for cognitive diagnosis that represents student proficiency as node-wise values on a concep…
2026-08-09 · cognitive-diagnosis, knowledge-tracing, student-modeling, generative-ai, adaptive-learning
🏷️ Adaptive Learning
> **Adaptive learning** — AI-driven educational systems that adjust content, pacing, and instructional strategies based on individual learner characteristics and performance. Adaptive learning is the …
🏷️ AI Education
> **AI Education** — the broad field encompassing both AI in education (using AI to teach) and AI literacy (teaching about AI). As the wiki's umbrella concept, AI education connects instructional tech…
🏷️ Automated Assessment
> **Automated assessment** — the use of AI to evaluate student work, from formative quizzes to high-stakes exams. Automated assessment spans multiple modalities — multiple-choice, short answer, essay,…
2026-08-09 · automated-grading, assessment-validity, formative-assessment, bias-mitigation, teacher-role
🏷️ Automated Essay Scoring
> **Automated Essay Scoring (AES)** — the use of AI to evaluate and score written essays, spanning traditional statistical approaches, fine-tuned language models, and increasingly accessible LLM-based…
🏷️ Benchmark
> **Benchmark** — standardized test suites and evaluation frameworks used to measure AI model performance on educational tasks. Benchmarks enable reproducible comparison across models and approaches, …
🏷️ Computational Thinking
> **Computational thinking** — a problem-solving approach involving decomposition, pattern recognition, abstraction, and algorithmic design. In AI education, computational thinking is both a prerequis…
🏷️ CS Education and AI
> **CS Education** — computer science education is the most-researched STEM subfield in the wiki, benefiting from natural alignment between AI tools and programming tasks. Code generation, debugging a…
2026-08-09 · computational-thinking, stem-education, automated-grading, prompt-engineering, higher-ed
🏷️ Generative AI
> **Generative AI** — AI systems capable of producing text, code, images, and other content, most prominently large language models like GPT-4 and Claude. Generative AI is the technology driving the c…
🏷️ Hallucination Risk
> **Hallucination Risk** — the danger that AI systems generate plausible but factually incorrect or fabricated content in educational contexts, where such errors can mislead learners, undermine trust,…
🏷️ Pedagogical Safety
> **Pedagogical safety** — the design principle that AI education systems must protect learners from harm, including inappropriate content, unsafe advice, biased treatment, and manipulative interactio…
🏷️ Professional Training and AI
> **Professional training** — the use of AI for workforce development, corporate learning, and professional skill acquisition. Professional training extends AI in education beyond formal schooling int…
🏷️ RAG (Retrieval-Augmented Generation)
> **RAG (Retrieval-Augmented Generation)** — an AI architecture that combines information retrieval with text generation, allowing LLMs to ground responses in external knowledge sources rather than re…
🏷️ Socratic Method
> **Socratic Method** — a pedagogical approach rooted in guided questioning and dialogue rather than direct instruction, now being adapted for generative AI tutoring systems. In AI in education, the S…
🏷️ Student Experience with AI
> **Student experience with AI** — how learners perceive, interact with, and are affected by AI tools in educational settings. With over 85 articles in the wiki, student experience is one of the most-…
📄 Generative AI and the Productivity Divide: Human-AI Complementarities in Education
> **Synthesis:** Idan & Anand (2026) conduct an RCT showing that GenAI access significantly increases task performance on average — but the gains are highly uneven, NOT predicted by GPA or prior knowl…
📄 Interactive learning dashboards: rethinking learning visualisations as engagement tools
> **Synthesis:** Graf et al. (2026) transformed a conventional Learning Analytics Dashboard (LAD) into an interactive ILAD by adding an LLM-powered pedagogical agent and a Judgement of Learning (JoL) …
2026-08-08 · learning-analytics, metacognition, higher-ed, engagement-metrics, self-regulated-learning
📄 When Help is Unhelpful: Evaluating AI Tutors for Productive Struggle
> **Synthesis:** Zhang et al. (2026) introduce TutorMoments, a replay-based evaluation framework that tests whether LM tutors adapt their pedagogical actions to context — scaffolding when support is n…
🏷️ Pedagogical Agent
> **Synthesis**: Pedagogical agents are AI-driven conversational interfaces embedded in learning environments that use pedagogical strategies (eliciting, telling, scaffolding) to support learner engag…
📄 Using LLMs to Detect Growth in Computational Thinking in Introductory Physics
> **Synthesis:** Savage, Shanker, Michlitsch & Rebello (2026) investigate using LLMs to evaluate students' written explanations of computational physics problems at scale. Establishing a human-coded b…
📄 What out-of-the-box LLMs can(t) do in law? A Turing test in Italian exams for lawyers, judges and notaries
> **Synthesis:** This paper reports a blind Turing Test evaluating leading LLMs on three Italian professional legal examinations: the Bar exam, Judges exam, and Notary exam. LLMs generated full writte…
📄 Learning to Use AI for Learning: Teaching Responsible Use of AI Chatbot to K-12 Students Through an AI Literacy Module
> **Synthesis:** An LLM-based interactive module teaches K-12 students prompting literacy through scenario-based deliberate practice with an AI auto-grader providing immediate, detailed feedback. Depl…
📄 WIP: Chat-Debugging: Large Language Model as a Hardware Debugging Assistant
> **Synthesis:** Work-in-progress exploring LLMs as debugging assistants for physical hardware lab courses. Proposes 'Chat-Debugging' where students interact with an LLM to diagnose circuit faults. Ai…
📄 AI-accelerated End-to-End Framework for Rapid Professional Upskilling
> **Synthesis:** The Crew Scaler framework applies AI acceleration across all five stages of professional upskilling—knowledge acquisition, content development, content review and verification, AI-tut…
📄 Instructional Agents: Reducing Teaching Faculty Workload through Multi-Agent Instructional Design
> **Synthesis:** Instructional Agents is a multi-agent LLM framework that automates end-to-end course material generation by simulating role-based collaboration among Teaching Faculty, Instructional D…
📄 Can LLMs Effectively Simulate Human Learners? Teachers' Insights from Tutoring LLM Students
> **Synthesis:** Semi-structured interviews with 12 teachers who tutored LLM-simulated students (MathDial dataset) reveal key authenticity gaps: overly complex language, lack of emotions, unnatural at…
🏷️ Help-Seeking
> **Help-Seeking** — a key concept in AI in education research. Explored across 4 articles in this wiki.…
📄 Calibrating Trustworthiness: Co-Designing Metrics and Visualizations for Evaluating LLMs in Education
> **Calibrating Trustworthiness: Co-Designing Metrics and Visualizations for Evaluating LLMs in Education** — Longitudinal co-design with learning engineers building an LLM-powered digital textbook. C…
📄 EduClaw-Bench: A Long-Horizon Benchmark for Pedagogical LLM Agents with Simulated Learners
> **EduClaw-Bench: A Long-Horizon Benchmark for Pedagogical LLM Agents with Simulated Learners** — Introduces a 30-day long-horizon benchmark for pedagogical LLM agents using simulated learners ground…
📄 CODE-GEN: A Human-in-the-Loop RAG-Based Agentic AI System for Multiple-Choice Question Generation
> **A dual-agent RAG-based system for generating and validating coding comprehension MCQs**, evaluated by 6 SMEs across 7 pedagogical dimensions (N=288 questions, 2,016 rating pairs). AI excels at cri…
📄 DeepTutor: Towards Agentic Personalized Tutoring
> **A fully open-source agentic tutoring framework that closes the loop between citation-grounded problem tutoring and difficulty-calibrated question generation**, powered by a hybrid personalization …
📄 EduZone: A Framework for Evaluating LLM Safety for K-12 Students and Teachers
> **EduZone is an automated evaluation framework that generates contextually grounded adversarial interactions to probe LLM safety in K-12 education, revealing that models are more vulnerable to educa…
📄 Bridging Artificial Intelligence and Power Systems Education Using a Hands-On Executable Framework
> **An open, executable module library for engineering-grounded AI (EGAI) in power systems education lowers the entry barrier for newcomers, with a progressive difficulty ladder from DNN templates to …
📄 Comparative Validation of GPT-4o-mini and Teacher Mean Scores for Automated Scoring of Music Analysis Responses: Single-Pass Deployment, Repeatability, and Strategy-Specific Bias
> **GPT-4o-mini can produce stable rubric-based scores for open-ended music analysis responses, with few-shot chain-of-thought prompting agreeing most strongly with teacher means while RAG systematica…
📄 From MOOC to MAIC: Reshaping Online Teaching and Learning through LLM-driven Agents
> **A new paradigm for online education replacing MOOCs with LLM-driven multi-agent AI classrooms**, piloted at Tsinghua University with 100K+ learning records from 500+ students. MAIC uses specialize…
📄 Agentic AI and Pedagogical Best Practice: The Tension Between Automation and Learning
> Education AI is shifting from passive chatbots to **proactive agents** that initiate and pursue goals. This offers personalisation but risks undermining **learner agency and cognitive effort**. The …
📄 Agreement Is Not Quality: Blind Expert Verification of Human and LLM Qualitative Coding When Human Consensus Is Not Ground Truth
> **Alex Liu, Lief Esbenshade, Michael Xiao, Victor Tian, Zachary Zhang, Kevin He, Min Sun** — arXiv preprint (2026).…
📄 From Planning to Revision: How AI Writing Support at Different Stages Alters Ownership
> Gero, Long, Schnitzler & Dhillon (2026, DIS '26) ran a between-subjects essay study (n = 253) showing that **where** AI support enters the writing process determines how much students feel they own …
2026-08-03 · writing-education, student-experience, ai-generated-content, metacognition, generative-ai
📄 From authentic products to authenticated processes: authentic assessment in AI-rich higher education
> Generative AI has not created the need for authentic assessment — it has made weaknesses in assessment design harder to ignore. Polished products can now be generated or substantially mediated by to…
📄 ConnectED: A Curriculum-Aligned AI System for Vietnamese Instructional Lesson Planning and Student Learning
> **Thang Doan Viet, Anh Nguyen Hoang, Tinh Luong Son, Anh Hoang Thi Ngoc, Huyen Giang Thi Thu, Tai Le Quy** — arXiv preprint (2026).…
📄 CoTAL: Human-in-the-Loop Prompt Engineering for Generalizable Formative Assessment Scoring and Feedback
> 1. **Evidence-Centered Design (ECD)** — assessments and rubrics aligned to curriculum goals from the start 2. **Human-in-the-loop prompt engineering** — labelled examples and prompts refined iterati…
2026-08-03 · formative-assessment, automated-grading, human-in-the-loop, prompt-engineering, benchmark
📄 Human-LLM Collaborative Inductive Coding for Conceptualizing K-12 Educator AI Use
> **Alex Liu, Min Sun, Lief Esbenshade, Michael Xiao, Victor Tian, Zachary Zhang, Kevin He** — arXiv preprint (2026).…
📄 Measuring Cognitive Engagement in Collaborative Discourse with an Extended ICAP Framework: Comparing Human Annotation, In-Context Learning, and Reflective LLM Agents
> **Lan Anh Do, Hanling Jiang, Shuchin Aeron, Ayanna K. Thomas** — CogSci 2026 (accepted full paper).…
📄 Let''s Chat: Leveraging Chatbot Outreach for Improved Course Performance
> Meyer, Page, Mata et al. (2026) ran two pre-registered RCTs at Georgia State University testing a **non-generative** academic chatbot that texted students 2–3 customized nudges per week in large-enr…
📄 To Facilitate or not to Facilitate: Human and LLM Facilitator Tendencies in Online Discussions
> **Dimitris Tsirmpas, Katerina Korre, John Pavlopoulos** — arXiv preprint (2026).…
📄 The Theoretical Foundation of Socratic Tests: Dynamic, Multimodal, Conversational Examinations
> **Ilya Mikhelson** — Submitted to Computers and Education: Artificial Intelligence (2026).…
📄 Advancing diagram-based reasoning in AI tutoring systems: a structural approach for STEM education
Presents **StructRAG**, a pattern-aware framework that improves how AI tutoring systems interpret **complex engineering diagrams** (circuit schematics, network topologies, block flowcharts) in STEM. C…
📄 Structured AI Demonstrations and Student LLM Use in Engineering Mechanics: Study Design and Preliminary Results
> **Shuang Geng, Helen Lallos-Harrell, Jiya Ashar, Thomas J. McKenna, Annwesa Dasgupta, Caleb Farny, Emma Lejeune** — arXiv preprint (2026).…
📄 Trust and Reliance on AI in Education: AI Literacy and Need for Cognition as Moderators
> Pitts, Rani & Mildort (2026, AIED) show with 432 undergraduates that **higher trust in an AI assistant is associated with lower appropriate reliance**: students who trusted the assistant more were w…
📄 Comparing human and LLM ordered coding of qualitative data: How coding differences cascade through temporal analysis
> **Authors:** Kamila Misiejuk, Sonsoles López-Pernas, Eduardo Araujo Oliveira, Brendan Eagan, Mohammed Saqr **Source:** Computers and Education: AI, Vol 11 — Open Access (CC BY 4.0)…
🏷️ Agentic AI in Education
> **Agentic AI** — AI systems that autonomously plan, execute, and adapt multi-step workflows to achieve learning goals, going beyond single-turn Q&A to act as persistent, goal-directed collaborators:…
🏷️ AI Tutoring
> **AI tutoring** — the use of AI (especially [[llm|LLMs]] and [[intelligent-tutoring|intelligent tutoring systems]]) to provide personalized, adaptive, scalable instructional support: conversational …
📄 Beyond Rephrasing: Book-Level Organization Improves Synthetic Textbook Data for Mid-Training
Studies how organizing synthetic content into coherent book-level documents affects language model training, moving beyond local rewriting. Presents a scalable synthesis pipeline that retrieves source…
📄 ICLE++: Modeling Fine-Grained Traits for Holistic Essay Scoring
Introduces ICLE++, a new annotated corpus of persuasive student essays that addresses critical limitations of the dominant ASAP benchmark in [[automated-essay-scoring]] research. Unlike ASAP — used by…
📄 IKS-Instruct: A 24,000-Example Multilingual Dataset for Teaching Language Models Indian Knowledge Systems
Presents a 24,795-example multilingual instruction dataset for teaching LLMs to deliver educational content grounded in Indian Knowledge Systems. Spans seven languages and bridges a gap in non-Western…
📄 A review of intervention designs of LLM Integration in Undergraduate Computer Science Education
This scoping review analyzed **13 experimental studies** on LLM integration in undergraduate [[cs-education]], examining how intervention design choices shape learning outcomes. The central finding: *…
📄 Rethinking LLM-Judged Helpfulness as a Pedagogy Signal: A Pre-Registered Audit Across Tutor Models
Pre-registered study auditing whether general-purpose helpfulness rubrics can distinguish direct answer-giving from pedagogical guidance in LLM tutors. Uses deterministic detectors for answer leakage …
📄 The Easy Trap: Why LLMs Underestimate Misconception-Driven Difficulty
LLMs systematically underestimate the difficulty of misconception-driven items ('The Easy Trap'). While LLM ratings show moderate rank correlation with empirical student difficulty (rho=0.52-0.70), th…
2026-07-30 · formative-assessment, adaptive-learning, feedback-loop, student-experience, stem-education
📄 Aligning LLM-Simulated and Human Examinees for Psychometric Calibration: A Cognitive Diagnostic Profiling Approach
Proposes Cognitive Diagnostic Profiling (CDP), a zero-shot framework that dramatically improves LLM-simulated examinee alignment with human test-takers. With CDP, IRT difficulty Spearman correlations …
📄 Agentic AI in Education: A Scoping Review of Research Landscape, Capabilities, and the Frontier Agent Paradigm
Published in *Computers and Education: Artificial Intelligence*, accepted 27 July 2026. 📄 doi:10.1016/j.caeai.2026.100653…
📄 SafeTutors: Pedagogical Safety in AI Tutoring
> **SafeTutors** is a benchmark that jointly evaluates safety and pedagogy in AI tutoring systems across mathematics, physics, and chemistry. It argues that **tutoring safety is fundamentally differen…
📄 Collaborative AI Literacy Framework
> **Collaborative AI Literacy Framework** — SEFI 2025. A systematic review of 9 studies (2015–2023) examining how collaborative learning (CL) approaches can be harnessed to build AI literacy across di…
📄 ISD Agent Benchmark
> **ISD-Agent-Bench** is a comprehensive benchmark for evaluating LLM-based instructional design agents, comprising **25,795 scenarios** generated via a Context Matrix framework that combines 51 conte…
📄 LLM Fallacy Misattribution in Education
> **The LLM Fallacy** is a cognitive attribution error in which individuals misinterpret LLM-assisted outputs as evidence of their own independent competence — producing a systematic gap (∆C) between …
📄 PersonaVLM: Long-Term Personalization for AI Tutors
> **PersonaVLM** introduces an agent framework for long-term personalization of multimodal LLMs, enabling AI tutors to remember, reason about, and align with a learner's evolving preferences across hu…
📄 Designing a mobile chatbot-based learning journaling system for intrinsic motivation and engagement
A **randomized 2×2 full-factorial field experiment** (N = 179 German university students, 22 days of app use, 12-week follow-up) testing two design principles for a **mobile chatbot-based learning jou…
2026-07-29 · self-regulated-learning, generative-ai, higher-ed, student-experience, engagement-metrics
📄 EduQwen: Pedagogical RL
> **EduQwen: Pedagogical RL** — A multi-stage optimization strategy combining reinforcement learning (DAPO) and supervised fine-tuning (SFT) to enhance the pedagogical knowledge of open-source LLMs, p…
📄 Multimodal Dialogue in STEM Education
> **The Multimodal Interference Effect** describes a systemic accuracy drop when LLMs encounter image-rich STEM problems: from ~96% on text-only physics problems to ~74% on multimodal ones. A simple t…
📄 Leveling the Playing Field: Temporal Video Segmentation for Individuals with ADHD in Computing Education
Pimenova, Begel and colleagues evaluate a post-hoc video processing intervention that segments instructional videos into single-instruction chunks with fixed pauses, reducing extraneous cognitive load…
📄 The Effect of High-Frequency, Automatically-marked Formative Assessments on Student Outcomes in A-Level Sciences
This quasi-experimental mixed-methods longitudinal study (N=142) deploys a fully automated marking pipeline for handwritten mock examinations in A-Level sciences, removing the human-marking bottleneck…
📄 A didactical-driven teacher assistant for a dimensional modeling course
Brisson, Segarra and Smits present a didactically-driven LLM teacher assistant for a university dimensional modeling (data warehousing) course. Unlike most educational chatbots that delegate pedagogic…
📄 Ordered Network Analysis of Epistemic Emotions during Collaborative Problem Solving
Anindho, Venkatesha, Ocumpaugh and Blanchard apply Ordered Network Analysis to trace how epistemic emotions such as confusion and frustration persist and transition during co-situated collaborative pr…
📄 Auditing Institutional Heterogeneity for Generative AI in Patient Education: A Large-Scale Study of 102 US Transplant Handbooks
Li, Padman and Krishnan audit 102 US transplant-center patient handbooks that serve as grounding corpora for generative AI patient-education assistants. They show large institutional heterogeneity in …
🏷️ Affective Computing
> **Affective computing** in education uses physiological and behavioral signals to sense learner emotion and adapt instruction — see [[affective-text-wearable-student-health]], [[multimodal-affective…
🏷️ Open Source
> **Open-source** AI in education is studied in [[lata-ferpa-compliant-local-llm-autograder]], [[vismatic-secure-sandbox-cs-education]], and [[open-source]] (tag) pages: local open models address [[pr…
🏷️ Prompt Engineering
> **Prompt engineering** — the practice of designing and refining inputs to large language models to achieve desired outputs. In education, prompt engineering serves dual roles: as a learner skill (st…
🏷️ Reinforcement Learning
> **Reinforcement learning** trains AI tutors and agents through reward signals: [[special-r1-rl-special-education]], [[singh-eduqwen-pedagogical-rl-2026]], [[pedagogical-safety-rl]], and [[ai-coachin…
2026-07-28 · pedagogical-safety, intelligent-tutoring, special-education, personalized-learning, k-12
📄 Beyond Perspectives: A Trio-Ethnography of Interpretation Evolution in LLM-Supported Programming Education
This experience report introduces trio-ethnography — structured dialogue between two computing educators with differing teaching philosophies and one undergraduate CS student — as a method for surfaci…
📄 Generative AI Availability, Grades, and Student Satisfaction at a Large University
This large-scale observational study tests the "GenAI substitution hypothesis" — the concern that students offload cognitive effort to [[generative-ai]] and earn inflated grades without learning. Usin…
📄 Distinguishing Artificial from Authentic: Evaluating LLMs for Detecting LLM-Generated Content
As students increasingly use [[llm]]s to draft written responses and program code, this study asks whether LLMs can reliably detect their own generated content across educational task types — programm…
📄 Exploring the Design Space of LLM-Based Programming Support in CS Education: A Scoping Review through the Lens of Assistance Governance
This scoping review synthesizes 90 peer-reviewed [[llm]]-based programming support systems in [[cs-education]] to make explicit how each system bounds, enacts, and controls assistance — decisions the …
📄 MedGame: Storytelling Gamification Empowered by Large Language Models for Medical Education
MedGame transforms static clinical cases into structured, executable storytelling games for medical education, moving beyond the localized question-answering and single-turn feedback that characterize…
📄 Representation Robustness under Executable Reasoning Constraints in Large Language Models for Mathematical Problem Solving
This study probes how sensitive [[llm]] mathematical problem solving is to the surface representation of an item — a question with direct bearing on [[assessment-validity]] when LLMs are used for scor…
📄 What Does the Credential Still Certify? Cognitive Stewardship for AI-Mediated Education
Generative AI undermines a basic premise of educational assessment: that submitted work reliably evidences the human capacities a credential certifies. This paper proposes *cognitive stewardship*, a f…
📄 Informal Learning Emerges in Everyday Human-LLM Interaction
As LLMs take over task execution, a central worry is that everyday AI use becomes cognitive offloading that erodes people's own capability development. This study analyses 128,569 naturalistic human-L…
📄 EduGuard: A Safe RAG-Based LLM Tutor for Programming Education
EduGuard is a retrieval-augmented generation (RAG) tutoring framework that directly confronts the safety and pedagogical failures of unrestricted LLM tutors in introductory programming. Unrestricted t…
📄 Student Evaluation of Repeated AI Feedback Across a Semester of Writing
This short paper provides rare descriptive classroom evidence on what happens when students repeatedly use generative-AI feedback across a full semester of writing coursework. Drawing on 2,988 reflect…
📄 Artificial intelligence and feedback in university education: effectiveness and student perceptions
This quasi-experimental study directly compares **AI-generated feedback** (two LLMs: **GPT-o4-mini** and **DeepSeek R1**) with **expert human-teacher feedback** in a project-based university course (A…
📄 Benchmarking Multimodal Large Language Models for Scientific Visualization Literacy
Multimodal large language models (MLLMs) are increasingly used to interpret visualizations, yet most evaluations remain chart-centric and offer limited insight into **scientific visualization (SciVis)…
📄 Learning Engagement Assistant (LEA): Cross-Course Scalability and Classroom Evaluation of an Agentic AI Tutoring System
LEA (Learning Engagement Assistant) is an **agentic AI tutoring system** that couples course-specific retrieval-augmented generation (RAG) with structured [[knowledge-tracing]] / Knowledge Component (…
📄 A Comparative Analysis of Institutional and Course Generative AI Policies within Higher Education: Implications for Instruction in Computing Education
> **Synthesis:** A comparative content analysis of institutional GenAI policies and computing-course syllabi in U.S. research-intensive universities, revealing a gap between broadly pro-use institutio…
📄 A Semi-Automated System for Generating Dialogue-Based TTS Lessons Using Large Language Models: An Exploratory Study of Educational Potential
**Gendo Kumoi, Fumie Watanabe, Tota Suko, Takashi Ishida, et al. (2026)** - arXiv preprint (IEEE). arXiv preprint. Kumoi, G., Watanabe, F., Suko, T., Ishida, T., et al. (2026). [A Semi-Automated Syste…
2026-07-15 · generative-ai, personalized-learning, scaffolding, active-learning, pedagogical-llm-training
📄 Commenting with Copilot: A Taxonomy and Multi-Year Analysis of Student Code-Generation Specifications
Analyzes how students specify intended behavior in natural language to AI code tools (Copilot) across multiple years, deriving a taxonomy of code-generation specifications expressed through comments. …
2026-07-14 · student-experience, stem-education, higher-ed, reshaping-cs-education-genai, ai-literacy
📄 Knowledge Distillation for Automated AI Tutor Evaluation
Addresses the lag between LLM integration into K-12/higher education and reliable methods for evaluating pedagogical quality. The authors introduce a knowledge-distillation approach to automate AI-tut…
📄 LLM-Generated Design Problems for Assessing Higher-Order Thinking in Project-Based Learning
Introduces 'design problems' (DPs): concise, scenario-based prompts that require applying knowledge in transfer contexts, generated with LLMs to assess higher-order thinking (HOT) in project-based lea…
📄 The Paternalistic Filter: Epistemic Injustice and Differential Refusal in LLM-Mediated History Education for Marginalized Romanian Students
A systematic API audit of four LLMs acting as history tutors evaluates 1,800 responses about the 1989 Romanian Revolution, exposing a 'paternalistic filter': models differentially refuse or soften ans…
📄 Q-Learning Lab: Teaching Reinforcement Learning Through Learner-Generated Trace Analysis
> Presents Q-Learning Lab, a single-file tool that makes the Bellman update concrete by letting undergraduates inspect how each value is computed and why actions are chosen, through learner-generated …
2026-07-14 · active-learning, higher-ed, reinforcement-learning, stem-education, self-regulated-learning
🏷️ Bias Mitigation
> **Bias mitigation** in educational AI requires auditing models across the pipeline: [[gender-bias-transfer-llm-writing]], [[ai-scoring-language-bias-physics]], [[llm-cultural-relevance-k12]], and [[…
📄 Using AI-based Learning Assistants in Higher Education: A Large-Scale Descriptive Analysis
Presents a large-scale descriptive analysis of an AI learning assistant (Syntea) using objective log data from 77,543 higher-education students, characterizing real usage patterns, adoption, and engag…
📄 From Execution to Education: A Bloom-Aligned Framework for Measuring Educational Control in LLMs
Introduces a Bloom-aligned framework for measuring 'educational control' in LLMs: the ability to preserve a task's instructional intent while shifting its cognitive demand toward higher-order Bloom le…
📄 How YouTube Frames ChatGPT Use in Education: An Epistemic Network Analysis with Supporting Multimodal Metadata
Uses epistemic network analysis of multimodal YouTube metadata (transcripts, titles, thumbnails, comments) to show how different creator groups frame ChatGPT use in education, revealing divergent narr…
📄 Why does AI unlock new possibilities in STEM education? A Bibliometric Analysis of Trends and Future Agenda
STEM education faces challenges in personalization and interdisciplinary integration. AI technology has brought new possibilities, but the mechanisms by which AI reshapes the STEM education ecosystem …
📄 Flowcode: An AI-Powered Programming Environment for Scaffolding Iteration in Creative Computing Education
Building upon found examples is a popular way people learn to code, especially in creative coding communities where sharing projects and remixing are common practices. But effectively doing so require…
📄 A Guiding Framework for K-12 Teachers in Creating AI-powered Learning Technologies through Vibe Coding
Large language models generate code from natural language prompts, enabling vibe coding, which allows non-programmers to develop computational solutions. Vibe coding for teachers amplifies the teacher…
📄 The GenAI Skill Bypass: Mapping Divergent Pathways of University Students and Staff AI Literacy
Higher education institutions are increasingly expected to ensure that both students and staff develop Generative AI (GenAI) literacies. In response, they are introducing professional development prog…
📄 Agents That Teach: Designing Incidental Learning Back into AI-Assisted Software Development
As AI coding agents take over substantial implementation work, developers increasingly lose the informal, effortful problem-solving through which software engineering expertise historically accumulate…
📄 CSTutorBench: Benchmarking Small Language Models as Tutors for Block-Based Programming
Deploying LLM tutors in K-12 raises concerns around privacy, cost, and reliance on proprietary models, motivating small language models (SLMs) as an alternative. The authors introduce **CSTutorBench**…
📄 Prompt Coach: An Empirical Evaluation of an Agentic Tutor for Learning Prompt Engineering in Software Development
Prompt engineering is a critical yet undertaught skill for software developers, poorly served by traditional instruction because of its evolving, interactive, context-dependent nature. The authors int…
📄 Say What? Examining Text and Voice Input Modalities for Prompt-Based Programming in Computing Education
Nearly all prior research on LLMs in computing education has used text input, yet voice-enabled interfaces are becoming common. This exploratory study investigated how introductory programming student…
📄 Understanding Student Perceptions, Mistakes, and Debugging Approaches when Solving Natural Language Programming Tasks
Learning to communicate with code-generating AI is an emerging skill for novice programmers. 'Prompt Problems' — having students solve computational tasks by writing natural-language prompts for code-…
📄 Automated Grading of Linux/Bash Examinations Using Large Language Models
**Manuel Alonso-Carracedo, Ruben Fernandez-Boullon, Pedro Celard, Francisco J. Rodriguez-Martinez, Lorena Otero-Cerdeira (2026)** This paper presents an [[llm]]-based grading system for Linux/bash com…
📄 Constructing Epistemic AI Literacy: Detecting Epistemic Aims and Processes in Student-AI Co-Programming
**Mengqian Wu (2026)** Epistemic thinking — understanding how knowledge is constructed and justified — plays a central role in [[ai-literacy]], particularly when students co-program with generative AI…
📄 Data Comics for Education: Evaluating Effectiveness, Benefits, and the Ethics of AI-Assisted Creation
Data comics combine sequential visual narratives with data visualization to improve student engagement with [[generative-ai]] in educational settings. This paper evaluates the effectiveness of AI-assi…
📄 Evaluating Interactivity: Toward Automated Assessment of AI-Generated Explorable Explanations
While [[llm]]s now enable rapid generation of learning materials like [[generative-ai]], evaluating the pedagogical quality of these materials remains an open challenge. This paper proposes an automat…
2026-07-03 · ai-generated-content, formative-assessment, learning-analytics, higher-ed, automated-grading
📄 From Answer Generators to Reasoning Facilitators: Designing AI Tutors for Mathematical Reasoning in High-Stakes Environments
The rapid integration of [[llm]]s into [[intelligent-tutoring]] threatens to reduce mathematical learning to mere answer generation. This paper presents a design framework for AI tutors that act as re…
📄 Mind the Trust Gap: Identifying (Mis)alignments in Teacher-Student Views Toward Control and Agency in K-12 Classroom AI
**Tomohiro Nagashima, Lisa Siegrist, Niklas Scholz, Shintaro Sato, Martina Vincoli, Man Su (2026)** As AI technologies enter [[k-12]] classrooms, understanding how different stakeholders perceive thes…
📄 Child Safety in Generative AI: An Expert-Guided and Incident-Grounded Evaluation Framework
> **Haein Kong** — HEAL Workshop at CHI 2026, submitted 1 Jul 2026…
📄 CogTax: A Four-Level Cognitive Taxonomy for Command-Line Computing Education
> **Manuel Alonso-Carracedo, Ruben Fernandez-Boullon, Pedro Celard, Francisco J. Rodriguez-Martinez, Lorena Otero-Cerdeira** — Universidade de Vigo, submitted 30 Jun 2026…
📄 Demystify, Use, Reflect, Assess (DURA): An Experience Report on LLM Integration in CS2
> **Margaret Ellis, Nikitha Donekal Chandrashekar, Sehrish Basir Nizamani, Mohammed Farghally, Jake O'Brien, Naren Ramakrishnan** — SIGCSE Virtual 2026, submitted 29 Jun 2026…
📄 ELEVATE: Designing Human-Centered GenAI Virtual Tutors for Scalable and Inclusive Education
> **Lorenzo Stacchio, Michele Giordano, Daniele Berardini, Primo Zingaretti, Emanuele Frontoni** — submitted 17 Jun 2026…
📄 Gaze-Informed Proactive AI Assistance for Children’s Picture Exploration
> **Zekun Wu, Man Su, Huiyong Li, Tomohiro Nagashima, Anna Maria Feit** — submitted 1 Jul 2026…
📄 Less Deliberate in Teams: Student LLM Use Across Individual and Collaborative Work
> **Sehrish Basir Nizamani, Zannah Ziew, Saad Nizamani, Khyati Goyal** — ACM SIGCSE Virtual 2026, submitted 29 Jun 2026…
2026-07-02 · student-experience, collaborative-learning, higher-ed, engagement-metrics, llm-in-education
📄 Visualizing Engineering Fundamentals: Design of Mixed Reality and Physical Toolkits for Effective Learning
> **Mohammad Abu Nasir Rakib, Sharmin Akter, Eshwara Prasad Sridhar, Somik Biswas, Md Rassel Raihan, Mahmudur Rahman** — submitted 1 Jul 2026…
📄 Touching and Feeling the Data: A Reusable Software Pipeline for Tactile Statistical Graphs in Accessible Education
> **Lawrence Obiuwevwi, Krzysztof J. Rechowicz, Jessica M. Johnson, Erika Frydenlund, Vikas Ashok, Sachin Shetty, Sampath Jayarathna** — IEEE IRI 2026, submitted 1 Jul 2026…
📄 Why Put in This Much Effort?": How AI Availability Shapes Students’ Motivation in Introductory Programming
**Tran, Harper & Price (2026)** examine a pressing motivational paradox in contemporary computing education: the ready availability of AI tools that can complete programming assignments undermines stu…
📄 AI in the Wild: A Large Scale Analysis of Authentic Interactions of College Students with Generative AI
Karidi, Amir & Roll (2026) present one of the largest empirical analyses to date of authentic (rather than lab-based) interactions between college students and generative AI tools. By analyzing intera…
📄 To Tab or Not to Tab: Measuring Critical Engagement in AI Code Completion Tools Using Behavioral Signals and Attention Checks
Hutchison et al. (2026) develop and validate a method for measuring critical engagement with AI code completion tools in educational settings. Using behavioral signals (time-to-accept, edit distance f…
📄 Invisible Impact of Empathy on Behavioral Change: Isolating the Effect of Empathy in Long-term Physical Activity Coaching Chatbot Interactions
Siyan et al. (2026) conduct a carefully controlled experiment isolating the effect of empathetic language in LLM-powered physical activity coaching chatbots over a longitudinal deployment. While the e…
📄 From Prompting to Epistemic Proactivity: Temporal Trajectories of Student-AI Interaction in Mathematics Learning
Abdelghani, Kaiser & Murayama (2026) trace how middle and high school students' interactions with AI math tutors evolve over time, identifying a trajectory from superficial prompting ('tell me the ans…
📄 Generative AI Literacy Training Improves Intelligence Analysts’ Discrimination of Real and AI-Generated Images
Kamali et al. (2026) evaluate a Generative AI Literacy training intervention designed to improve intelligence analysts' ability to distinguish real photographs from AI-generated images. In a controlle…
📄 Exploring the Value of Diverse LLM Explanations in Introductory Programming
Bernstein, Denny, Leinonen et al. (2026) investigate whether providing students with multiple, diverse LLM-generated explanations of code (rather than a single 'best' explanation) improves comprehensi…
📄 Four Types of LLM Reliance and Their Predictors Among Undergraduate Writers: A Mixed-Methods Study at a Minority-Serving R1 University
Hossain (2026) develops a typology of LLM reliance among undergraduate writers at a minority-serving R1 institution, identifying four distinct profiles: strategic scaffolders who use AI for idea gener…
📄 Teaching Prompt-Based Programming with LLMs: A 45-Minute Lesson with Guided Practice for End-User Programmers
This study by Tran, Marwan & Price (2026) introduces and evaluates a 45-minute structured lesson on prompt-based programming, a new modality enabled by LLMs where users express computational goals thr…
📄 A Survey of Automated Presentation Coaching: Systems, Methods, and Open Challenges
This survey provides the first systematic review of automated presentation coaching systems, organizing them along a five-dimensional task taxonomy: segmental pronunciation, lexical stress, suprasegme…
📄 DysLexLens: A Low-Resource LLM Framework for Analysing Dyslexic Learners Insights from Online Forums
DysLexLens is a low-resource LLM framework designed to analyze how [[special-education|dyslexic learners]] experience AI tools by mining online forum discussions. The framework employs dictionary-driv…
📄 Cognitive Episodes in LLM Reasoning Traces Enable Interpretable Human Item Difficulty Prediction
This paper introduces Epi2Diff (Episode to Difficulty), a framework that maps LLM reasoning traces into cognitively grounded episode sequences for predicting human item difficulty in [[assessment|educ…
📄 An exploratory behavioral and electroencephalographic study of artificial intelligence-assisted learning modes in high school students
📄 [PDF](https://arxiv.org/pdf/2606.26579) This study investigates how different modes of AI interaction affect cognitive engagement and learning outcomes in high school students. Using a within-subje…
📄 The impact of generative artificial intelligence on academic development of Chinese students in humanities and social sciences
This large-scale survey of humanities and social sciences (HSS) students in China examines how [[generative-ai]] reshapes academic development across four dimensions: usage patterns, effects on learni…
📄 SupplyNet: Supporting Visual Exploratory Learning in Supply Chain via Contextual Multi-Agent Simulation
SupplyNet is a gamified visual simulation system that uses a contextual graph-based [[llm]] multi-agent framework to model interdependent supply chain dynamics. Designed for [[professional-training]] …
📄 What Changes When the Interlocutor Is an AI? Interactional Fluency and Linguistic Uptake in L2 Spoken Dialogue
Scheinberg et al. (2026) analyze 78 university learners of German across four sites completing a counterbalanced spot-the-difference task with both a human peer and a real-time AI partner. Using diari…
📄 Analysis and Prediction of At-Risk Students Using Machine Learning Algorithms
Gheisari and Salarian (2026) apply supervised machine learning classification to identify at-risk students before they withdraw from higher education programs. The study evaluates Logistic Regression,…
📄 WIP: Bridging the Gap Between Instructional Design and Pedagogical Use: A Framework for Mathematics Educators
Castillo Ventura et al. (2026) address the gap between instructional design of digital mathematics resources and their pedagogical use in classrooms. Their work-in-progress framework translates learni…
📄 The Correct Answer Trap: Pedagogically-Grounded Detection and Feedback for Hidden Misconceptions
Imran and Bulathwela (2026) identify the 'correct answer trap' — automated feedback systems that judge only answer correctness reinforce rather than address misconceptions when students reach the righ…
📄 CourseBlueprint: A Structured Pipeline for Adaptive Pedagogical Video Generation Grounded in Course Corpora
Islam et al. (2026) address a core limitation of generative text-to-video for education: while visually fluent, such systems lack pedagogical content knowledge (PCK). CourseBlueprint provides a struct…
📄 Curiosity as Linguistic Intervention: Using LLM Tutoring Dialogues to Influence Exploratory Learning Behavior
Ganganath et al. (2026) introduce CURIOBOT, a framework that operationalizes Berlyne's four collative variables (novelty, complexity, conflict, uncertainty) as adaptive linguistic interventions in con…
2026-06-23 · intelligent-tutoring, metacognition, scaffolding, active-learning, self-regulated-learning
📄 Why Machines Misread Pedagogical Quality: Human-Machine Alignment in LLM-Based Pretest Question Evaluation
Tseng et al. (2026) investigate human-machine alignment in LLM-based pretest question evaluation — a critical bottleneck for scalable AI-assisted assessment. Their AI-assisted workflow combines automa…
📄 Test-Driven, AI-Assisted Learning: Replacing Lectures with Weekly Closed-Book Tests
Liu et al. (2026) report on a 13-week Test-Driven, AI-Assisted (TDAA) redesign of a Theory of Computation course at HKUST (Guangzhou). The course replaced all lectures with self-directed, AI-assisted …
🏷️ Knowledge Tracing
> **Knowledge tracing** — modeling what learners know over time by tracking their performance on exercises and predicting future mastery. It is the wiki's richest modeling thread, spanning Bayesian, d…
📄 Learning to Prompt: Improving Student Engagement with Adaptive LLM-based High-School Tutoring
> **Po-Chin Chang, Nicholas Hogan, Aske Plaat, Michiel T. van der Meer** (2026). arXiv cs.AI preprint…
📄 PsyScore: A Psychometrically-Aware Framework for Trait-Adaptive Essay Scoring and ZPD-Scaffolded Feedback
> **Wei Xia, Jin Wu, Haoran Shi, Xiangyu Wang, Chanjin Zheng** (2026). East China Normal University / arXiv cs.CL preprint…
📄 Toward Accessible Psychotherapy Training Using AI-Driven Interactive Patient Avatars
AI-driven interactive patient avatars for psychotherapy training provide accessible, repeatable practice with measurable skill improvement in evidence-based therapy techniques. Toward Accessible Psych…
📄 ASTRA: A Scalable Next-Generation ATCO Training Simulator with Autonomous Simpilots
ASTRA uses autonomous AI sim-pilots for scalable air traffic control training; reduces dependency on human role-players while maintaining realistic scenario complexity. ASTRA: A Scalable Next-Generati…
📄 Engagement Intensity as a Learner-Modeling Signal for Adaptive AI Ethics Instruction
> Engagement intensity during AI ethics instruction serves as an effective learner-modeling signal for adaptive instruction; prior LLM experience influences engagement patterns.…
📄 From Memorization to Creation: Evaluating the Cognitive Depth of LLM-Generated Educational Questions
LLM-generated educational questions show varying cognitive depth; models excel at factual recall but struggle with higher-order thinking questions per Bloom's taxonomy. From Memorization to Creation: …
📄 Contaminated Collaboration: Measuring Gender Bias Transfer in LLM-Assisted Student Writing
> **Ariyan Hossain, Kazi Kamruzzaman Rabbi, Farig Sadeque, S M Taiabul Haque** (2026). arXiv cs.CL…
📄 LecturaAgents: A Multi-Agent Framework for Adaptive Personalized AI-Assisted Learning and Embodied Teaching
> **Jaward Sesay, Yue Yu, Siwei Dong, Yemin Shi, Guangyao Chen, Borje F. Karlsson** (2026). arXiv cs.CL…
📄 ParaTutor: LLM Mediated Parent Child Tutoring through Role Separated Scaffolding Interface in Real Time
> **Lan Luo, Anqi Wang, Muzhi Zhou, Junhua Zhu, Jie Cai, Ao Yu, Hui Pan** (2026). arXiv cs.HC…
📄 Rethinking Scaffolding in LLM Tutors: The Interactional Mismatch Between Benchmarks and Real-World Deployments
> **Alexandra Neagu, Jeffrey T. H. Wong, Marcus Messer, Rhodri Nelson, Peter B. Johnson** (2026). Pluralistic Alignment Workshop @ ICML 2026…
📄 Using AI in engineering education: a balancing act, driven by clear purpose
Based on a questionnaire of 100 higher-education engineering students and a critical literature review, examines how students use and perceive LLMs. Students value LLMs for writing support, conceptual…
📄 AI as a Partner in Learning about, Doing, and Engaging with Science: Vigilance as the Key to Productive Augmentation
Argues that epistemic vigilance — the human evaluation of AI output calibrated to how far a fallible source can be trusted — is the binding constraint on productive augmentation. AI's fluent, confiden…
📄 Cross-Dataset Bloom Question Classification: Supervised Models and Prompted LLMs
Evaluates cross-dataset generalization of ML/DL methods and LLMs for automatic Bloom's taxonomy classification of assessment questions across five datasets. Supervised ML/DL models degraded substantia…
📄 Improving Capstone Team Outcomes through Dynamic Skill Matching and Preference Alignment
Team-based projects are a cornerstone of engineering and computing courses, but unstructured team formation often leads to poor project outcomes due to misaligned student interests and inadequate skil…
2026-06-16 · intelligent-tutoring, edtech-platform, higher-ed, stem-education, personalized-learning
📄 The Missing Layer: Why EdTech Needs Design-Time Generative UI, Not Just Runtime Personalization
> Argues the dominant paradigm of runtime GenUI adaptation in EdTech is insufficient. Proposes design-time card-based GenUI where educational content is encoded as modality-agnostic semantic units and…
📄 Gender Differences in AI Literacy Workshop Outcomes and Deepfake Engagement
> Examines gender differences in AI literacy, safety awareness, and STEM career aspirations among Australian secondary students (Years 7, 8, 10; N=199) from two co-educational government schools after…
📄 What do you mean by human-AI collaboration: Prerequisite functions and the affordances needed to achieve it
> Asks what is gained and lost when 'collaboration' is applied freely to human-AI interaction. Argues true collaboration requires symmetric/negotiated relationship, shared goals, low and shifting divi…
📄 LearnOpt: Recovering the Latent Cognitive Structure of Standardized Examinations via Knowledge Graphs and Constrained Optimization
Standardized examinations are typically treated as uniform syllabus coverage problems. LearnOpt recovers stable latent cognitive structures diverging systematically from official syllabi, using LLM-ta…
📄 Are LLM-based Chatbots Good Enough to Support Computer Science Students in Multiple-Choice Exercises?
Investigates LLM chatbots' performance on 70 MCQs for a university CS lecture on interactive visual data analysis, comparing with student performance. GPT-4o and GPT-5 significantly outperformed small…
📄 Measuring Whether LLM Tutors Teach or Solve: A Diagnostic for Educational Impact
Studies whether public LLM tutoring benchmarks distinguish learning-supportive behavior from mere answer production. Proposes a lightweight diagnostic based on the gap between solving-oriented and ped…
📄 Leveraging Physiological Signals to Predict Exam Outcomes with Machine Learning
> Investigates ML models to predict exam outcomes from physiological data (electrodermal activity, heart rate, skin temperature) collected during exams. Evaluates logistic regression, random forest, S…
📄 Stuck in a Spiral": Shame and Guilt as Social Regulators of AI Use in Computing Education
> An interview study with 19 computing students through a functionalist perspective of shame and guilt. Findings show these emotions regulate when and how students make their AI use visible, engaging …
📄 Simulating Students' Java Programming Errors with Large Language Models
This paper investigates whether [[llm|large language models]] can serve as scalable proxies for students by simulating realistic logical errors in code submissions. Using the CodeWorkout dataset of 74…
2026-06-15 · stem-education, student-experience, intelligent-tutoring, learning-analytics, efficacy-study
📄 Exploring How Agent Voice Accents Shape Human-AI Collaboration in K-12 Group Learning
**Ravi, Stevens, Hurt, Hanks, Lin & Anderson (2026)**. Ravi et al. investigate how the voice accent of a [[generative-ai]] conversational peer agent shapes learners' perceptions, trust, and interactio…
📄 AiAWE: An Open-Source LLM Automated Writing Evaluation System Using LoRA-Adapted Instruction-Tuned Models
Gayed presents **AiAWE**, an open-source [[automated-grading|automated writing evaluation]] (AWE) system that scores argumentative essays using a LoRA-adapted instruction-tuned [[llm|large language mo…
📄 Generativism: Toward a Learning Theory for the Age of Generative Artificial Intelligence
**Li & Zheng (2026)**. Li & Zheng argue that the four dominant learning theories — behaviorism, cognitivism, constructivism, and connectivism — show significant conceptual limitations as [[generative-…
📄 The Environmental Cost of LLMs in AIED: Reporting and Practices
> **Sabrina C. Eimler, Lukas Erle, Daniel Flood, Aditi Haiman, Luca Häckert, André Helgert, Lachlan McGinness, Büsra Yapici**…
📄 Learning by Chatting? Investigating the Impact of Generative AI on Information Seeking and Learning
> **Shravika Mittal, Su Lin Blodgett, Q. Vera Liao**…
📄 The Empirically Grounded Adaptive Virtual Patient for Psychotherapy Training
**Angela Chen, Siwei Jin, Catherine Bao, Canwen Wang, Robert E. Kraut, Tongshuang Wu, Haiyi Zhu** — cs.CY, cs.HC The Adaptive Virtual Patient (AVP) is an LLM-driven simulated patient for psychotherapy…
2026-06-10 · professional-training, generative-ai, intelligent-tutoring, student-experience, higher-ed
📄 AI-Integrated Learning Management System for Middle School: A Longitudinal Study of Learning Outcomes
**Misan Paul Etchie, Taiwo Olutosin** — cs.CY, cs.AI, cs.HC This paper proposes an AI-integrated LMS designed specifically for middle school instruction, addressing the gap between current LMS platfor…
2026-06-10 · k-12, adaptive-learning, personalized-learning, formative-assessment, intelligent-tutoring
📄 AI-Driven Analytics of Team-Teaching Talk: Acoustic Patterns across Experience, Cohorts and the Learning Design
**Yuchen Liu, Roberto Martinez-Maldonado, Riordan Alfredo, Paola Mejia-Domenzain, Dwi Rahayu, Sadia Nawaz** — AIED 2026 — cs.HC, cs.AI This paper presents an AI-based speech processing approach to ana…
📄 Profiling cognitive offloading in LLM-mediated synthesis writing: Volume vs. content
**Oleksandra Poquet, Mani Shankar Nanduri, Maria Ximena Salinas Loyer, Matthias Stadler, Michael Sailer, Jelena Jovanovic** — Accepted at EC-TEL 2026 — cs.HC, cs.ET This study compares two approaches …
📄 Reexamining the Cold-Start Problem in Knowledge Tracing Models and Implications for SafeInsights
**Jiayi Zhang, Ryan S. Baker, Debshila Basu Mallick, Cristina Heffernan, Neil Heffernan** — cs.HC This paper replicates and extends prior work on the cold-start problem in knowledge tracing — the chal…
📄 EduMirror: Modeling Educational Social Dynamics with Value-driven Multi-agent Simulation
**Jingzhe Lin, Hengbin Yu, Yongdan Zeng, Fangwei Zhong** — ICML 2026 — cs.MA, cs.CY EduMirror introduces a multi-agent simulator for studying educational social dynamics, addressing the dilemma that o…
📄 Report on CHIIR 2026 Workshop on Generative AI and Academic Search (GAI&AS)
**Yifan Liu, Jaime Arguello, Orland Hoeber, Chang Liu et al.** — cs.IR, cs.AI, cs.HC This report summarizes the CHIIR 2026 Workshop on Generative AI and Academic Search (GAI&AS), which examined how Ge…
📄 Hybrid E-Assessment in Higher Education: Semi-Automated Grading of Paper-Based Written Examinations
**Hartwig Grabowski, Michael Canz** — cs.AI, cs.CV, cs.CY This paper identifies the didactic narrowing caused by fully digital e-assessment (overuse of closed question formats) and proposes a hybrid a…
📄 Detecting Knowledge Gaps from Conversational AI Interactions Using Curriculum Prerequisite Graphs
This paper introduces a pipeline that maps student questions directed at a conversational AI teaching assistant to curriculum topics using a few-shot text classifier, grounded in a GPT-4-extracted pre…
📄 Design and Implementation of a Real-time Multi-site Immersive Learning System Using Photon Fusion
> This paper develops a VR-based immersive learning environment using Photon Fusion that allows teachers and students to be present in the same virtual space regardless of physical locations. The syst…
📄 Reshaping Undergraduate Computer Science Education in the Generative AI Era
**Yi-Chieh Lee, Nattapat Boonprakong, Yugin Tan, Harold Soh et al.** — Workshop report from NUS-Google Workshops — cs.CY This white paper synthesizes findings from two international NUS-Google Worksho…
📄 TibetCPR: A Multimodal Tactile Feedback System for CPR Training in High-Altitude Regions
**Yibo Meng, Ruiqi Chen, Zhiming Liu, Xiaolan Ding** — Accepted at MobileHCI 2026 — cs.HC TibetCPR is a low-cost, self-guided CPR training system that pairs depth-driven electrotactile feedback with r…
2026-06-10 · professional-training, formative-assessment, student-experience, edtech-platform, higher-ed
📄 FOXGLOVE: Comparing Goal-Oriented Writing Feedback from Experts and LLMs
Introduces **FOXGLOVE**, a dataset of 696 feedback comments by trained writing instructors on 69 twelfth-grade argumentative essays, paired with 1,644 comments from four frontier LLMs — totaling 2,340…
📄 Role of Instructional Guidance in Generative AI-Assisted Learning
Investigates how instructional guidance shapes student-AI interaction in [[higher-ed|construction engineering education]]. Introduces a **five-step prompting framework** grounded in Generative Learnin…
📄 LLM-Generated Feedback in Introductory Programming: A Classroom Study
Presents a **large-scale classroom study** (N=215 students, 6,693 submissions across 17 labs) deploying AI-generated feedback through a randomized protocol in an introductory Python programming course…
📄 Regulating the AI Tutor: SRL and Help-Seeking in Adolescent GenAI Use
Examines how 98 Grade-9 students across three German Gymnasium schools regulated their use of a Mistral-Large GenAI tutor while preparing for a math exam. Despite overwhelmingly selecting scaffolded s…
📄 AI Assistance for Discretionary Work: Increasing Feedback Provision in Higher Education
This field experiment shows that AI-generated feedback drafts can measurably increase the rate and length of feedback that teaching assistants actually deliver to students, without sacrificing perceiv…
📄 Teacher-Authored Prompts for Configuring Student-AI Dialogue: K-12 Classroom Implementation
This large-scale K-12 deployment provides empirical evidence that teacher-authored prompts can reliably shape the cognitive quality of student-AI dialogue at classroom scale. The TASD system lets teac…
📄 Warning About AI Fallibility Increases Help-Seeking in an Intelligent Tutoring System
> **Synthesis:** Recent work in Technology-Enhanced Learning and HumanComputer Interaction highlights the importance of transparency and trust calibration in AI-supported learning environments as they…
📄 AI-Generated Traces for Novice Programmers: Learning Effects and Learner Differences in a Multi-Institutional Study
Multi-institutional study on Generated Animated Traces (GATs) for CS1. Found that mid-engagement students may experience a performance decrement due to coordination costs (Expertise-Reversal Effect). …
📄 Conversational AI as a catalyst for informal learning: An empirical large-scale study on LLM use in everyday learning
> **Synthesis:** Conversational AI as a catalyst for informal learning: An empirical large-scale study on LLM use in everyday learning…
📄 Balancing AI responsibility with privacy, safety, and utility: Unlearning in large language models for mathematics education
> **Synthesis:** Balancing AI responsibility with privacy, safety, and utility: Unlearning in large language models for mathematics education…
📄 VETTING: A dual-LLM framework for in-loop safety verification via policy isolation in educational AI
> **Synthesis:** VETTING: A dual-LLM framework for in-loop safety verification via policy isolation in educational AI…
📄 Beyond Access: Guided LLM Scaffolding for Independent Learning in Undergraduate Statistics
> Experimental study comparing Guided vs. Unrestricted LLM access. Explicit training in reasoning-focused scaffolding (stepwise hints, verification) led to significantly better independent performance…
📄 Special-R1: Reinforcement Learning for Special Education — Aligning LLM Tutors to Diverse Learners through Disability-Adaptive Training
> **Authors:** Unggi Lee, Jihoi Na, Yeil Jeong, Haeun Park, Yeonju Jang (2026)…
2026-06-01 · intelligent-tutoring, special-education, personalized-learning, reinforcement-learning, k-12
📄 The Tutoring Effectiveness Index: Predicting LLM Math Tutor Quality from Four Conversation Signals
> **Authors:** Shim Jaechang, Unggi Lee (2026) — CIKM 2026…
2026-06-01 · intelligent-tutoring, benchmark, efficacy-study, automated-grading, formative-assessment
📄 Benchmarking Large Language Models for Diagnosing Students' Cognitive Skills from Handwritten Math Work
> **MathCog** benchmark (3,036 teacher-annotated diagnostic verdicts, 639 handwritten responses, 18 LLMs): all models severely underperform (macro F1 < 0.5) — over-attributing evidence, overthinking m…
📄 Generalizing a Highly Configurable Analytics Pipeline to Replicate and Support Educational Research Across Multiple Domains
Artificial intelligence assistants deployed in online learning environments create new opportunities to collect large volumes of learner interaction data and generate insights to improve student outco…
📄 Who Am I? History-Aware Profiles for Student Simulation in Tutoring Dialogues
A key part of developing large language model (LLM)-powered, automated tutoring tools is student simulation, i.e., using LLMs to role-play as students, which can facilitate tutor model evaluation and …
📄 Modularizing Educational LLM-Agency for Fostering Responsible Learning Assistance
The widespread adoption of AI chatbots in education will drastically change learning, making responsible deployment a critical concern. While large language models (LLMs) might have access to sources …
📄 It's OK Because...": The Wild West of Student Rationalization of AI Use in Academic Writing
Generative AI challenges academic integrity not only by enabling students to delegate substantial portions of their academic work, but also by blurring the ethical boundaries by which students disting…
🏷️ AI Ed Evaluation
> **AI-ed evaluation** — the body of methods, benchmarks, and criteria used to assess whether AI education tools (LLM-based tutors, automated graders, feedback systems, agents) actually work — not jus…
🏷️ AI Plagiarism Detection
Technologies and methods for detecting AI-generated content in academic submissions, including classifier-based approaches, watermarking, and stylistic analysis. The effectiveness and reliability of t…
📄 Agentic Literacy Debt: A Structural Problem the AI Literacy Field Has Not Yet Named
**Agentic Literacy Debt** names a critical gap in the [[ai-literacy]] landscape that has become urgent with the rise of autonomous AI agents. Existing AI literacy frameworks assume humans evaluate AI …
📄 Smaller, Younger, and More Impactful: How AI-Assisted Writing Transforms Research Teams
**AI-Assisted Writing Transforms Research Teams** challenges the longstanding "Big Science" trend toward ever-larger teams, showing that AI writing tools enable smaller, younger research teams to prod…
📄 Catching The Correct Answer Trap: Characterising AI Tutor Blind Spots When Analysing Student Reasoning
**Catching the Correct Answer Trap** — accepted at AIED 2026 — exposes a critical blind spot in [[intelligent-tutoring]] systems: they systematically fail to detect misconceptions when students arrive…
2026-05-28 · intelligent-tutoring, automated-grading, formative-assessment, scaffolding, generative-ai
📄 Mathematical Modelling of Ethical AI Use in Higher Education: A Coordination Game Framework for Future-Facing Learning
**Ethical AI Use in Higher Education: A Coordination Game Framework** provides a formal mechanism-level account of why policy statements alone fail to change student AI-use behavior. Reframing student…
📄 KT4EQG: Personalized Exercise Question Generation via Knowledge Tracing
**KT4EQG: Personalized Exercise Question Generation via Knowledge Tracing** bridges two key AI-in-education paradigms: [[personalized-learning]] through question generation and [[learning-analytics]] …
📄 LLM-assisted sentiment analysis for integrated computational and qualitative mixed methods education research: A case study of students' written reflection assignments
**LLM-Assisted Sentiment Analysis for Mixed-Methods Education Research** demonstrates how LLMs can serve as scalable qualitative research assistants, enabling researchers to investigate multiple demog…
2026-05-28 · edtech-platform, higher-ed, learning-analytics, student-experience, formative-assessment
📄 Learning after COVID-19 and the ICT career aspirations: Are students entering the AI era with weaker skills?
**Post-COVID ICT Career Aspirations** uses PISA 2018 and 2022 country-level data to investigate whether students entering the generative AI era have adequate educational foundations. Using a mixed-met…
📄 REC-CBM: Rubric-Aware Error-Correction Concept Bottleneck Models for Trustworthy Open-Ended Grading
**REC-CBM: Rubric-Aware Error-Correction Concept Bottleneck Models** advances the [[automated-grading]] frontier by solving a fundamental trust problem: even accurate AI graders are unusable if educat…
📄 Generative AI and the marginalization of minoritized knowledges in higher education: the case of disability
This paper argues that [[generative-ai]] systems in [[higher-ed]] are not epistemically neutral — they actively marginalize non-hegemonic ways of knowing. Drawing on educational sciences, critical tec…
📄 Persistent AI Agents in Academic Research: A Single-Investigator Implementation Case Study
This is the first empirical study of what happens when AI agents are embedded **persistently** in a real academic research environment — with durable memory, local files, external tools, scheduled rou…
📄 Slide Deck Q&A Quality Assurance App: A Multi-Stage Pipeline for Pedagogical Question Generation
SlidesQAQA is a Flask-based system that extracts text and rendered images from PDF lecture slides and processes them through a four-stage [[llm]] pipeline: **window planning** (segment extraction), **…
📄 How Students (Mis)understand Conditionals and Loops -- A Taxonomy
This paper presents a fine-grained taxonomy categorizing novice programmers' difficulties with reading and understanding control flow constructs — specifically conditionals (selection) and loops (iter…
📄 Codify: An Intelligent Socratic Tutoring System for Programming Education
📄 DOI: 10.32473/flairs.39.1.141554 Codify (also called AI Tutor) is an [[intelligent-tutoring]] system that leverages [[llm|LLMs]], competency tracking, and adaptive assessment to provide Socratic, d…
2026-05-26 · intelligent-tutoring, stem-education, higher-ed, adaptive-learning, self-regulated-learning
📄 Generative AI as a Design Variable: An Evidence-Centered Framework for Principled Governance in STEM Assessment
This paper proposes a principled framework grounded in Evidence-Centered Design (ECD) that treats [[generative-ai]] as a design variable within STEM assessment arguments rather than an external threat…
📄 It Felt a Bit Eerie": Exploring Humanlike Interactions During Collaborative Writing with an Artificial Agent
This comparative user study (n=48) examines how the temporal and visual dimensions of AI collaboration shape the experience of [[writing-education|writing tasks]], revealing that humanlike design feat…
📄 Explaining Too Much? Understanding How Large Language Model Reasoning Traces Influence Performance and Metacognition
This preregistered between-subjects study (N=559) provides the first rigorous evidence that [[llm]] reasoning traces — increasingly common in AI interfaces — do not improve performance and can activel…
2026-05-26 · metacognition, student-experience, efficacy-study, over-reliance, self-regulated-learning
📄 A Taxonomy of Metacognitive Learning Scenarios in Professional Contexts: Integrating Systems Theory with Empirical Constraints
This paper addresses a fundamental gap in [[metacognition]] research: the lack of systematic integration of metacognitive theories into scenario taxonomies capable of guiding AI-enhanced professional …
2026-05-26 · metacognition, professional-training, adaptive-learning, lifelong-learning, scaffolding
📄 Cognitive offloading and the speedup illusion in human-AI interaction
This preregistered large-scale study (N = 1,237) investigates whether people are well-calibrated in estimating the time savings from AI assistance on simple cognitive tasks. The key finding is a **spe…
📄 MindCopilot: Towards Formalizing and Evaluating Granular Human-LLM Co-Writing
MindCopilot introduces a formal framework for evaluating human-LLM co-writing that shifts from output-only metrics (BLEU, ROUGE) to **interaction-aware evaluation**. The paper models co-writing as a *…
📄 Socially fluent AI decouples conversational signals from source identity in online interaction
This study embedded undisclosed AI agents as teammates in synchronous text-based group interactions across analytical, creative, and ethical tasks with 786 participants making 1,572 identity judgments…
📄 Exploring the Effectiveness of Using LLMs for Automated Assessment of Student Self Explanations in Programming Education
This paper presents a rigorous empirical comparison between [[llm|LLM]]-based and semantic similarity methods for [[automated-grading|automated assessment]] of student self-explanations in programming…
📄 AI-Enabled Serious Games: Integrating Intelligence and Adaptivity in Training Systems
Serious games are widely used for learning and training across domains such as healthcare, defense, and education. This chapter examines how contemporary AI approaches may support real-time instructio…
📄 Expert Cognition Dashboard: From Learning Analytics to Cognition Intelligence in AI-Driven Education
**Annie Yuan (2026)**. arXiv preprint (cs.HC). Current AI-driven educational systems primarily rely on behavioural analytics and performance metrics, lacking the ability to model expert cognition used…
📄 Automated Grading of Handwritten Mathematics Using Vision-Capable LLMs
Automated grading systems have enabled scalable assessment for many response types, but handwritten mathematics remains a barrier due to the complexity of multi-step solutions. Vision-capable large la…
📄 Simulating Learners' Task-Selection Strategies and System Constraints in Mastery Learning
Intelligent Tutoring Systems often grant learners shared control over skill and problem selection. We propose a simulation-based framework to examine how learner task-selection strategies and system c…
2026-05-22 · intelligent-tutoring, mastery-learning, adaptive-learning, engagement-metrics, simulation
📄 ANVIL: Analogies and Videos for Lecturers
Noviello, Birillo, and Migut (2026) present ANVIL, an end-to-end multimodal generation pipeline for educational content — one of the first systems to automate the full journey from concept definition …
📄 Creating Learning Scaffolds for Engineering Design Using Concept Catalyst
Singh, Mansi, and Riedl (2026) present Concept Catalyst, an LLM-powered tool designed to reduce K-12 teacher preparation time for Engineering Design Challenges. Unlike general-purpose chatbots, Concep…
📄 What Makes Words Hard? Sakura at BEA 2026 Shared Task on Vocabulary Difficulty Prediction
🔗 [Code](https://github.com/adno/vocabulary-difficulty) This paper presents two complementary approaches to predicting vocabulary difficulty for language learners, achieving state-of-the-art results …
📄 PromptDecipher: Supporting AI Tutor Authoring Through Editable Simulated Interactions
Teachers virtually never test AI tutoring bots before student deployment; PromptDecipher enforces QA as a first-class activity by letting teachers edit bot responses directly. PromptDecipher addresses…
📄 CLARA: An AI-Augmented Analytics Dashboard for Collaboration Literacy
Agentic analytics using AI-produced concept-map artifacts as shared human-AI representations improves collaboration quality analysis and AI response grounding over transcript-only baselines. CLARA int…
📄 The Hidden Cost of Contextual Sycophancy: an AI Literacy Intervention in Human-AI Collaboration
LLM sycophancy creates a feedback loop where user errors propagate into AI advice, degrading outcomes; AI literacy training reduces but doesn't eliminate this contextual sycophantic dependence. This A…
📄 Generative AI Feedback, English Writing and Teacher Rubrics: A Multiple-Case Study of CyberScholar
RAG-based rubric-grounded GenAI writing feedback improved student revision quality (N=143, grades 7-11) and saved teacher time, but automated ratings were inconsistent. CyberScholar demonstrates rubri…
📄 Confirming Correct, Missing the Rest: LLM Tutoring Agents Struggle Where Feedback Matters Most
LLM tutors achieve near-ceiling on correct steps but systematically over-reject valid-suboptimal reasoning and over-validate incorrect solutions — precisely where adaptive tutoring matters most. This …
📄 An Interpretable Closed-Loop Intelligent Tutoring System for Multimodal Affective Feedback in Asynchronous Presentation Training
Closed-loop ITS with multimodal affective scoring (facial, vocal, textual, oculomotor) produced significant presentation skill gains (Cohen's d = 0.39-0.90, N=204) over 30 days. This paper presents on…
2026-05-19 · intelligent-tutoring, affective-computing, multimodal, higher-ed, professional-training
📄 Towards SocratiCode: Designing a Generative AI-Based Programming Tutor for K-12 Students through a 4-Week Participatory Design Study
Socratic questioning, reflection prompts, misconception checks, and mandatory pauses produce better K-12 engagement than directive answer-giving AI tutors. SocratiCode demonstrates a participatory des…
📄 The Effects of Structured LLM-Generated Feedback on Programming Assignment Performance
LLM-generated feedback produces faster time-to-solution than compiler-only baseline; counterintuitively, less guided feedback showed stronger effects than more guided variants. This study provides emp…
📄 A Formative Study of Brief Affective Text as a Complement to Wearable Sensing for Longitudinal Student Health Monitoring
In a year-long study of 458 university students (3,610 person-waves) using Oura rings for passive physiological sensing, researchers examined whether **ultra-brief affective text prompts** (median 3-w…
2026-05-17 · affective-computing, student-experience, higher-ed, learning-analytics, affective-tutoring
📄 Codify: An Intelligent Socratic Tutoring System for Programming Education
Codify (also referred to as "AI Tutor") is a web-based [[intelligent-tutoring]] platform for programming education that integrates conversational AI, adaptive assessment, and learning analytics. It le…
📄 AI-Driven Tools for Enhancing Campus Well-being: Prevention and Intervention
This dissertation presents an integrated AI framework for campus well-being spanning prevention (improving feedback collection) and intervention (advancing mental health detection). It represents an i…
📄 What Don't You Understand? Using Large Language Models to Identify and Characterize Student Misconceptions About Challenging Topics
This paper presents a systematic two-stage methodology for surfacing student misconceptions at scale. Drawing on 3,802 medical student enrollments across 5 biomedical science courses (9 course periods…
2026-05-16 · generative-ai, formative-assessment, higher-ed, learning-analytics, personalized-learning
📄 Simulating Students or Sycophantic Problem Solving? On Misconception Faithfulness of LLM Simulators
This paper exposes a critical failure mode in using LLMs as simulated students for [[intelligent-tutoring]] development and evaluation. The authors introduce **misconception faithfulness** — the prope…
📄 Retrieval-Augmented Tutoring for Algorithm Tracing and Problem-Solving in AI Education
KITE (Knowledge-Informed Tutoring Engine) introduces a [[intelligent-tutoring]] architecture that grounds its responses in course materials through a multimodal [[scaffolding|RAG pipeline]]. Unlike ge…
📄 Are Agents Ready to Teach? A Multi-Stage Benchmark for Real-World Teaching Workflows
> Are Agents Ready to Teach? A Multi-Stage Benchmark for Real-World Teaching Workflows **Chen et al. (2026)** — Multiple institutions. Under review.…
📄 Sycophancy is an Educational Safety Risk: Why LLM Tutors Need Sycophancy Benchmarks
> Sycophancy is an Educational Safety Risk: Why LLM Tutors Need Sycophancy Benchmarks **Kasneci & Kasneci (2026)** — Position paper. arXiv cs.AI/cs.HC.…
📄 Understanding How International Students in the U.S. Are Using Conversational AI to Support Cross-Cultural Adaptation
> Understanding How International Students in the U.S. Are Using Conversational AI to Support Cross-Cultural Adaptation **Nourian et al. (2026)** — Multiple institutions. arXiv cs.HC.…
📄 LaTA: A Drop-in, FERPA-Compliant Local-LLM Autograder for Upper-Division STEM Coursework
> LaTA: A Drop-in, FERPA-Compliant Local-LLM Autograder for Upper-Division STEM Coursework **Rodríguez (2026)** — Oregon State University. Submitted to Computers & Education.…
📄 LearnMate^2: Design and Evaluation of an LLM-powered Personalized and Adaptive Support System for Online Learning
> LearnMate^2: Design and Evaluation of an LLM-powered Personalized and Adaptive Support System for Online Learning **Wang, Lee, & Mutlu (2026)** — University of Wisconsin-Madison. CHI-related publica…
📄 Evaluating Prompt Injection Defenses for Educational LLM Tutors: Security-Usability-Latency Trade-offs
> Evaluating Prompt Injection Defenses for Educational LLM Tutors: Security-Usability-Latency Trade-offs **Maiorano (2026)** — arXiv cs.CR/cs.AI.…
2026-05-15 · intelligent-tutoring, generative-ai, regulation, hallucination-risk, student-experience
📄 Characterizing Students' LLM Usage Behaviors and Their Association with Learning in Critical Thinking Tasks
> Characterizing Students' LLM Usage Behaviors and Their Association with Learning in Critical Thinking Tasks **Park, Orozco Vasquez, & Conati (2026)** — University of British Columbia. Accepted at ED…
📄 Taklif.AI: LLM-Powered Platform for Interest-Based Personalized College Assignments
> Taklif.AI: LLM-Powered Platform for Interest-Based Personalized College Assignments **Kurdya et al. (2026)** — Multiple institutions. arXiv cs.AI.…
📄 AI-Generated Slides: Are They Good? Can Students Tell?
This study evaluated five generative AI tools for creating instructional slides from instructor-authored course notes: NotebookLM, Claude, M365 Copilot, Cursor, and Claude Code. Educators assessed sli…
📄 AICoFe: Implementation and Deployment of an AI-Based Collaborative Feedback System for Higher Education
AICoFe orchestrates a multi-LLM pipeline using GPT-4.1-mini, Gemini 2.5 Flash, and Llama 3.1 to synthesize quantitative rubric data and qualitative observations into actionable feedback for higher edu…
📄 Distinguishing performance gains from learning when using generative AI
This *Nature Reviews Psychology* piece draws a critical distinction that has been under-theorized in AIED research: The authors argue that generative AI easily boosts performance but often bypasses th…
📄 Children's English Reading Story Generation via Supervised Fine-Tuning of Compact LLMs with Controllable Difficulty and Safety
Using an expert-designed children's reading curriculum and stories generated by GPT-4o and Llama 3.3 70B as training data, the authors fine-tuned three different 8B-parameter LLMs. **The fine-tuned 8B…
📄 Interpretable Difficulty-Aware Knowledge Tracing in Tutor-Student Dialogues
This paper bridges LLM-based dialogue tutoring and interpretable student modeling. By mapping opaque LLM representations to **Item Response Theory** parameters — student ability (θ) and question diffi…
📄 Reinforcement Learning Measurement Model
Interactive assessments generate sequential process data that conventional item response models (IRT) cannot adequately handle. This paper proposes a **reinforcement learning measurement model** that …
📄 AcademiClaw: When Students Set Challenges for AI Agents
> **Yu, Lu, Si et al. (77 authors, 2026)** — Shanghai Jiao Tong University, SII, GAIR. Open-source benchmark.…
📄 When AI Agents Teach Each Other: Discourse Patterns Resembling Peer Learning in the Moltbook Community
> **Authors:** Eason Chen, Ce Guan, A Elshafiey, Zhonghao Zhao, Joshua Zekeri, Afeez Edeifo Shaibu, Emmanuel Osadebe Prince **Year:** 2026 **Venue:** arXiv (cs.HC) > Mining discourse from Moltbook, a …
2026-05-11 · agentic-ai, benchmark, collaborative-ai-tutoring, engagement-metrics, learning-analytics
📄 Cognitive Agent Compilation for Explicit Problem Solver Modeling
**Cognitive Agent Compilation (CAC)** is a framework that uses a strong teacher LLM to compile problem-solving knowledge into an explicit, inspectable target agent. Unlike end-to-end LLM tutoring appr…
📄 The Path to Conversational AI Tutors: Integrating Tutoring Best Practices and Targeted Technologies to Produce Scalable AI Agents
> **Authors:** Kirk Vanacore, Ryan S. Baker, Avery H. Closser, Jeremy Roschelle **Year:** 2026 **Venue:** arXiv (cs.HC) > Synthesizes intelligent tutoring systems research and generative AI into a kee…
📄 Not All Students Engage Alike: Multi-Institution Patterns in GenAI Tutor Use
> **Authors:** Youjie Chen, Xixi Shi, Xinyu Liu, Shuaiguo Wang, Tracy Xiao Liu, Dragan Gašević **Year:** 2026 **Venue:** arXiv (cs.CY) > Large-scale analysis (N=11,406 students, 200 classes, 10 instit…
📄 Modernizing Ground Truth: Four Shifts Toward Improving Reliability and Validity in AI in Education
> Stop treating κ > 0.8 as a binary stamp of approval.…
📄 Beyond the AI Tutor: Social Learning with LLM Agents
Most AI-based educational tools adopt a one-on-one tutoring paradigm, pairing a single LLM with a single learner. Yet decades of learning science — from Vygotsky's Zone of Proximal Development to Band…
📄 LLM-based Multimodal AI Feedback Produces Equivalent Learning and Better Student Perceptions than Educator Feedback
**AI multimodal feedback matches educator feedback for learning while significantly outperforming it on student perceptions.** The authors built a real-time AI-facilitated multimodal feedback system i…
📄 TeachingCoach: A Fine-Tuned Scaffolding Chatbot for Instructional Guidance to Instructors
> **Authors:** Isabel Molnar, Peiyu Li, Si Chen, Sugana Chawla, James Lang, Ronald Metoyer, Ting Hua, Nitesh V. Chawla **Year:** 2026 **Venue:** arXiv (cs.AI) > **Year:** 2026 > **Venue:** arXiv (cs.A…
📄 Higher Education Must Bridge the AI Gap
> A Science editorial by University of Illinois Chicago Chancellor Marie Lynn Miranda (April 2026) arguing that higher education has a narrow window to shape AI's distribution equitably. Proposes a th…
📄 Building AI Companions that Prioritise Learning over Performance
> A design framework for LLM-powered educational agents that prioritize durable learning over short-term task performance. Introduced by Khosravi et al. (2026), AI learning companions are defined as a…
📄 The Missing Evaluation Axis: What 10,000 Student Submissions Reveal About AI Tutor Effectiveness
> A framework for evaluating AI tutoring systems that extends beyond pedagogical quality of feedback to measure what students actually *do* with that feedback — whether they act on it and whether they…
📄 AISSA: AI-based Student Slides Analysis Tool for Academic Presentations
> A web-based system that uses LLMs and Learning Analytics dashboards to provide automated, rubric-based feedback on student presentation slides. Developed by Becerra et al. (2026), AISSA addresses th…
2026-05-09 · automated-grading, learning-analytics, formative-assessment, higher-ed, human-in-the-loop-ai
📄 A New Direction for Students in an AI World: Prosper, Prepare, Protect
> A yearlong global "premortem" by the Brookings Center for Universal Education (2026) examining generative AI's risks and benefits for students. Based on 500+ interviews across 50 countries, 400+ stu…
📄 The Pedagogy of AI Mistakes: Fostering Higher-Order Thinking
> An instructional approach that deliberately leverages AI errors, hallucinations, and limitations as teaching tools to foster higher-order thinking. Rather than viewing AI mistakes as failures to be …
📄 Prober.ai: Gated Inquiry-Based Feedback via LLM-Constrained Personas for Argumentative Writing
> A web-based writing environment that inverts the AI-tutoring paradigm: rather than generating improved text for students, Prober.ai constrains an LLM to ask only targeted inquiry-based questions abo…
🏷️ AI from the Administrator Perspective
> Stub — pending source ingestion. AI adoption, strategy, and governance from the institutional administrator and leadership perspective.…
🏷️ Lifelong Learning and AI
> Stub — pending source ingestion. Lifelong learning and AI support for continuous education beyond formal schooling.…
📄 AI Literacy Assessment: Self-Reported vs Performance Misalignment
>Highlights critical misalignment between self-reported AI literacy and actual performance. Teachers overestimate their AI skills by 40% on average. Performance-based assessments correlate better (r=0…
📄 ECNUClaw: A Learner-Profiled Intelligent Study Companion Framework for K-12 Personalized Education
> ECNUClaw is an open-source framework by Zhou, Li & Zhang (2026) for building **learner-profiled intelligent study companions** in K-12 education. The system maintains a **five-dimension learner prof…
📄 Generate-Then-Validate: Question Generation for Education
> **Synthesis:** A novel generate-then-validate pipeline for educational question generation that reduces LLM hallucination by 62% compared to direct generation, validated on STEM datasets with 89% ac…
📄 LLMs for Culturally Relevant K-12 Pedagogy
> Explores LLMs to support K-12 teachers in designing culturally relevant pedagogy. An exploratory pilot with four K-12 teachers found the CulturAIEd tool enhanced teachers' confidence in identifying …
📄 NSMQ Riddles: A Benchmark of Scientific and Mathematical Riddles for Quizzing Large Language Models
> Boateng et al. (2026) introduce **NSMQ Riddles**, a benchmark of 1.8K scientific and mathematical riddles drawn from 11 years of Ghana's **National Science and Maths Quiz** — a live TV competition f…
📄 Pedagogical Safety in Educational Reinforcement Learning
> As reinforcement learning personalizes instruction in intelligent tutoring systems, there is no formal framework for pedagogical safety — a critical gap. > First formal framework for defining and de…
📄 Programming Intelligent Tutoring Systems
> **SCRIPT** (Deriyeva, Dannath, Paassen, 2026) implements an intelligent tutoring system for **Python programming** in a German university context, filling a gap in prior ITS which rarely supported P…
2026-05-08 · intelligent-tutoring, stem-education, higher-ed, adaptive-learning, formative-assessment
📄 Quality-Conditioned Agreement in Automated Short Answer Scoring: Mid-Range Degradation and the Impact of Task-Specific Adaptation
> Schleifer, Ariely & Klebanov (2026) investigate a critical gap in [[automated-grading]]: **how scoring quality degrades for mid-range student responses**. Most ASAS evaluations focus on clearly corr…
2026-05-08 · automated-grading, formative-assessment, benchmark, efficacy-study, human-in-the-loop-ai
📄 TeachBench - Evaluating LLM Teaching Ability
> While LLMs are increasingly used as teaching assistants, their teaching capability remains insufficiently evaluated — a critical gap in current AIED research. > Syllabus-grounded framework for measu…
🏷️ Automated Question Generation
Automated question generation leverages NLP and LLMs to create educational assessments at scale. Wei & Stamper (2025) introduced the **generate-then-validate** paradigm, reducing hallucination by 62% …
🏷️ Culturally Relevant Pedagogy
Culturally Relevant Pedagogy (CRP), introduced by Gloria Ladson-Billings (1995), centers marginalized students' cultural references in curriculum design. Wang et al. (2025) demonstrate that **LLMs can…
🏷️ K-12 AI Education
K-12 AI Education encompasses the integration of artificial intelligence literacy, tools, and pedagogical approaches into primary and secondary education. Recent research reveals three critical pillar…
🏷️ Teacher AI Competency
Teacher AI competency encompasses the knowledge, skills, and dispositions required for effective AI integration in educational contexts. Emerging frameworks identify three competency dimensions: Zhang…
📄 AI Peer Feedback Systems
> Peer feedback develops critical reflection and evaluative judgment, yet: > Student peer feedback is often superficial or inconsistent. **AICoFe** (AI-based Collaborative Feedback) uses a multi-LLM p…
📄 AI Tutor Safety and Pedagogical Harms
> Conventional LLM safety benchmarks focus on toxic outputs, jailbreaks, and bias. In education, the primary risks are quieter: > "Solving problems correctly and avoiding toxic language does not make …
📄 Automatic Short Answer Grading with LLMs
> Automatic Short Answer Grading (ASAG) is never perfect. Upper bounds on accuracy arise from: > Zero-shot LLMs perform strongly on ASAG without task-specific fine-tuning, but **model-based confidence…
📄 Educational LLM Alignment
> Hardy & Kim (2026) identify a **cascading proxy** problem in AI-for-education evaluation: > The gap between what LLMs are *capable* of and what actually *benefits learners* — benchmark performance, …
📄 Educational VLM Evaluation
> Benchmarking vision-language models (VLMs) not on their ability to solve problems, but on their ability to *support learners* — particularly struggling learners and those making errors. Traditional …
📄 Interpretable Knowledge Tracing via IRT
> Two critical gaps in dialogue-based Knowledge Tracing (KT): > Most LLM-based dialogue tutoring systems produce opaque predictions. Huang et al. map raw LLM logits into **student ability (θ)** and **…
2026-05-07 · adaptive-learning, intelligent-tutoring, personalized-learning, learning-analytics, k-12
📄 LLM-Based Educational Simulation: Evaluating Temporal Student Persona Stability Across ADHD Profiles
> Gonnermann-Müller, Haase & Leins (2026) evaluate whether **LLM-generated student personas simulating ADHD profiles** maintain stable and realistic behavioral patterns over time. This addresses a cri…
📄 The LLM Fallacy and Misattribution of Competence
> Three system properties enable the fallacy via two cognitive mediators: > The LLM fallacy is a **cognitive attribution error** in which users misinterpret LLM-assisted outputs as evidence of their o…
📄 LLM Student Modeling and Long-Term Memory Architecture
> Current AI tutoring systems treat each session as independent. Adaptive systems use real-time knowledge tracing (e.g., [[knowledge-tracing-irt|IRT-based models]]) but rarely retain a longitudinal st…
📄 From Surface Learning to Deep Understanding: A Grounded AI Tutoring System for Moodle
> Ostrowska, Kukla & Majstrak (2026) present an AI tutoring system **integrated into the Moodle LMS** designed to scaffold students from surface-level fact recall to deep conceptual understanding thro…
2026-05-07 · intelligent-tutoring, higher-ed, edtech-platform, scaffolding, adaptive-learning-systems
📄 Multimodal AI Tutoring in STEM
> When LLMs process STEM problems that require interpreting diagrams, graphs, or schematics alongside text, their accuracy degrades substantially. This effect is: > General-purpose LLMs achieve near-c…
📄 Tutoring-Specific vs. General-Purpose AI in Education
> 1. **Desirable difficulties** — General-purpose AI removes productive struggle; tutoring tools preserve it via graduated hints. 2. **Germane load** — Effective learning requires processing that feel…
2026-05-07 · intelligent-tutoring, generative-ai, personalized-learning, scaffolding, adaptive-learning
🏷️ Affective Tutoring
> Integrating emotional awareness into AI tutoring systems can yield measurable pedagogical gains, but the same affective sophistication risks amplifying harms if learner agency is eroded by empatheti…
🏷️ AI Literacy
> **AI literacy** — the knowledge, skills, and critical dispositions needed to understand, evaluate, and effectively use AI technologies in educational contexts. AI literacy spans foundational underst…
🏷️ Formative Assessment in AI Education
Assessment designed to inform ongoing instruction and learning, as opposed to summative evaluation. AI systems can generate, validate, and adapt formative assessment items at scale, though quality var…
🏷️ Human-in-the-Loop AI for Education
Educational AI systems that strategically interleave automated generation with human expert judgment, preserving pedagogical quality while scaling production. Two recent implementations illustrate dis…
🏷️ Metacognition
> Metacognition — thinking about one's own thinking — is both a target of AI education research (can AI tools develop students' metacognitive skills?) and a risk factor (AI completing tasks may suppre…
🏷️ Training Pedagogical LLMs for Tutoring
> Domain-specialized optimization can transform a mid-sized open-source model (Qwen3-32B) into a pedagogical domain expert that outperforms far larger proprietary systems — but only when training rewa…
🏷️ Personalized Learning
Tailoring educational experiences to individual learner profiles, including prior knowledge, learning pace, preferences, and affective states. AI enables personalization at scale, though the gap betwe…
2026-05-07 · personalized-learning, intelligent-tutoring, adaptive-learning, ai-education, higher-ed
🏷️ Self-Regulated Learning
> Self-regulated learning (SRL) describes learners as active participants who can shape and develop their cognitive and behavioral actions in a successful way. AI tools can either scaffold SRL develop…
🏷️ Socratic AI Dialogue
> Socratic dialogue — asking structured questions rather than providing answers — is one of the strongest pedagogical scaffolds for deep learning. When automated via AI, it produces measurable reasoni…
📄 A meta-analysis of the effect of generative AI on productivity and learning in programming
> Maier, Gunzenhäuser & Schweisthal (2026) conduct a **meta-analysis synthesizing evidence** on how generative AI tools affect both programming productivity and learning outcomes. This is a **confiden…
📄 Towards Self-Referential Analytic Assessment: A Profile-Based Approach to L2 Writing Evaluation with LLMs
> Bannò, Knill & Gales (2026) propose a paradigm shift in automated essay scoring: from **inter-learner ranking** to **intra-learner profiling**. Instead of asking "how does this essay rank against ot…