🧠 AI Ed Wiki

🏷️ benchmark

54 pages tagged with benchmark(49 articles, 5 concepts)

📄 ELBench: A Multi-Dimensional Benchmark for Education-Facing Large Language Models
> **Synthesis:** Jiang et al. (2026) introduce **ELBench**, the first benchmark to evaluate education-facing LLMs on all four required dimensions — General Capability, Safety and Trustworthiness, Basi…
2026-08-13 · llm-evaluation, ai-ed-evaluation, pedagogical-safety, llm, safety
📄 Findings of the First Teaching Monster Challenge: A Benchmark of Pedagogical Content Knowledge in AI Agents
> **Synthesis:** Lin et al. (2026) present the **Teaching Monster Challenge**, the first instructional-video generation benchmark that treats the learner persona as an explicit evaluation criterion, m…
2026-08-13 · llm-evaluation, ai-ed-evaluation, agentic-ai, pedagogical-agent, content-generation
🏷️ Research Methods in AIED
> **Research methods in AIED** — the set of empirical designs, data-collection strategies, and analytic techniques researchers use to study AI in education: whether and how AI tools support (or harm) …
2026-08-13 · ai-education, educational-measurement, efficacy-study, rct, methodology
📄 Adversarial Stress Testing of Role-Playing Language Agents using Multi-Agent Evaluation
> **Synthesis:** This paper presents a modular multi-agent platform for adversarially stress-testing [[agentic-ai|role-playing language agents]] through structured multi-turn dialogue. With three coor…
2026-08-09 · agentic-ai, multi-agent, safety, evaluation, llm
🏷️ Benchmark
> **Benchmark** — standardized test suites and evaluation frameworks used to measure AI model performance on educational tasks. Benchmarks enable reproducible comparison across models and approaches, …
2026-08-09 · ai-ed-evaluation, assessment, llm, generative-ai
📄 When Help is Unhelpful: Evaluating AI Tutors for Productive Struggle
> **Synthesis:** Zhang et al. (2026) introduce TutorMoments, a replay-based evaluation framework that tests whether LM tutors adapt their pedagogical actions to context — scaffolding when support is n…
2026-08-08 · intelligent-tutoring, scaffolding, llm, llm-evaluation, k-12
📄 What out-of-the-box LLMs can(t) do in law? A Turing test in Italian exams for lawyers, judges and notaries
> **Synthesis:** This paper reports a blind Turing Test evaluating leading LLMs on three Italian professional legal examinations: the Bar exam, Judges exam, and Notary exam. LLMs generated full writte…
2026-08-07 · llm, assessment, professional-training, ai-ed-evaluation, automated-grading
📄 Robust and Efficient Motion Reasoning for Privacy-Aware Classroom Incident Recognition
> **Synthesis:** Pilot study on privacy-aware computer vision for classroom incident detection. Introduces a hybrid benchmark combining generative CCTV-style videos with real classroom pose data. Prop…
2026-08-06 · k-12, privacy, multimodal, classroom, ai-detection
📄 When Agents Learn to Be You: Benchmarking Privacy Leakage, Impersonation Risk, and Defenses in Persona Skills
> **When Agents Learn to Be You: Benchmarking Privacy Leakage, Impersonation Risk, and Defenses in Persona Skills** — Introduces AntiSkillBench with 7,500 persona-grounded dialogue traces from 50 beha…
2026-08-05 · privacy, agentic-ai, student-ai-interaction, bias-mitigation, personalized-learning
📄 EduClaw-Bench: A Long-Horizon Benchmark for Pedagogical LLM Agents with Simulated Learners
> **EduClaw-Bench: A Long-Horizon Benchmark for Pedagogical LLM Agents with Simulated Learners** — Introduces a 30-day long-horizon benchmark for pedagogical LLM agents using simulated learners ground…
2026-08-05 · intelligent-tutoring, llm, agentic-ai, knowledge-tracing, student-modeling
📄 EduZone: A Framework for Evaluating LLM Safety for K-12 Students and Teachers
> **EduZone is an automated evaluation framework that generates contextually grounded adversarial interactions to probe LLM safety in K-12 education, revealing that models are more vulnerable to educa…
2026-08-04 · llm, k-12, pedagogical-safety, ai-tutor-safety-harms, ai-governance-education
📄 CoTAL: Human-in-the-Loop Prompt Engineering for Generalizable Formative Assessment Scoring and Feedback
> 1. **Evidence-Centered Design (ECD)** — assessments and rubrics aligned to curriculum goals from the start 2. **Human-in-the-loop prompt engineering** — labelled examples and prompts refined iterati…
2026-08-03 · formative-assessment, automated-grading, human-in-the-loop, prompt-engineering, ai-ed-evaluation
📄 ICLE++: Modeling Fine-Grained Traits for Holistic Essay Scoring
Introduces ICLE++, a new annotated corpus of persuasive student essays that addresses critical limitations of the dominant ASAP benchmark in [[automated-essay-scoring]] research. Unlike ASAP — used by…
2026-07-31 · automated-essay-scoring, automated-grading, educational-measurement, formative-assessment, higher-ed
📄 Rethinking LLM-Judged Helpfulness as a Pedagogy Signal: A Pre-Registered Audit Across Tutor Models
Pre-registered study auditing whether general-purpose helpfulness rubrics can distinguish direct answer-giving from pedagogical guidance in LLM tutors. Uses deterministic detectors for answer leakage …
2026-07-31 · llm, intelligent-tutoring, automated-grading, feedback-loop, adaptive-learning
📄 Aligning LLM-Simulated and Human Examinees for Psychometric Calibration: A Cognitive Diagnostic Profiling Approach
Proposes Cognitive Diagnostic Profiling (CDP), a zero-shot framework that dramatically improves LLM-simulated examinee alignment with human test-takers. With CDP, IRT difficulty Spearman correlations …
2026-07-30 · llm, formative-assessment, adaptive-learning, student-experience, higher-ed
📄 ISD Agent Benchmark
> **ISD-Agent-Bench** is a comprehensive benchmark for evaluating LLM-based instructional design agents, comprising **25,795 scenarios** generated via a Context Matrix framework that combines 51 conte…
2026-07-29 · agentic-ai, ai-ed-evaluation, rag, llm, agentic-workflows
📄 MedGame: Storytelling Gamification Empowered by Large Language Models for Medical Education
MedGame transforms static clinical cases into structured, executable storytelling games for medical education, moving beyond the localized question-answering and single-turn feedback that characterize…
2026-07-24 · llm, generative-ai, professional-training, engagement-metrics, ai-tutoring
📄 Representation Robustness under Executable Reasoning Constraints in Large Language Models for Mathematical Problem Solving
This study probes how sensitive [[llm]] mathematical problem solving is to the surface representation of an item — a question with direct bearing on [[assessment-validity]] when LLMs are used for scor…
2026-07-24 · llm, stem-education, assessment-validity, reinforcement-learning, rag
📄 EduGuard: A Safe RAG-Based LLM Tutor for Programming Education
EduGuard is a retrieval-augmented generation (RAG) tutoring framework that directly confronts the safety and pedagogical failures of unrestricted LLM tutors in introductory programming. Unrestricted t…
2026-07-20 · llm, generative-ai, intelligent-tutoring, stem-education, over-reliance
📄 Benchmarking Multimodal Large Language Models for Scientific Visualization Literacy
Multimodal large language models (MLLMs) are increasingly used to interpret visualizations, yet most evaluations remain chart-centric and offer limited insight into **scientific visualization (SciVis)…
2026-07-17 · llm, generative-ai, stem-education, ai-literacy, higher-ed
📄 Learning Engagement Assistant (LEA): Cross-Course Scalability and Classroom Evaluation of an Agentic AI Tutoring System
LEA (Learning Engagement Assistant) is an **agentic AI tutoring system** that couples course-specific retrieval-augmented generation (RAG) with structured [[knowledge-tracing]] / Knowledge Component (…
2026-07-16 · llm, generative-ai, intelligent-tutoring, higher-ed, stem-education
📄 CSTutorBench: Benchmarking Small Language Models as Tutors for Block-Based Programming
Deploying LLM tutors in K-12 raises concerns around privacy, cost, and reliance on proprietary models, motivating small language models (SLMs) as an alternative. The authors introduce **CSTutorBench**…
2026-07-08 · llm, intelligent-tutoring, k-12, privacy, cs-education
📄 Cognitive Episodes in LLM Reasoning Traces Enable Interpretable Human Item Difficulty Prediction
This paper introduces Epi2Diff (Episode to Difficulty), a framework that maps LLM reasoning traces into cognitively grounded episode sequences for predicting human item difficulty in [[assessment|educ…
2026-06-29 · assessment, llm, learning-analytics, higher-ed, student-modeling
📄 Rethinking Scaffolding in LLM Tutors: The Interactional Mismatch Between Benchmarks and Real-World Deployments
> **Alexandra Neagu, Jeffrey T. H. Wong, Marcus Messer, Rhodri Nelson, Peter B. Johnson** (2026). Pluralistic Alignment Workshop @ ICML 2026…
2026-06-17 · scaffolding, intelligent-tutoring, llm, efficacy-study, student-experience
📄 Measuring Whether LLM Tutors Teach or Solve: A Diagnostic for Educational Impact
Studies whether public LLM tutoring benchmarks distinguish learning-supportive behavior from mere answer production. Proposes a lightweight diagnostic based on the gap between solving-oriented and ped…
2026-06-16 · intelligent-tutoring, llm, feedback-loop, scaffolding, learning-analytics
📄 Reexamining the Cold-Start Problem in Knowledge Tracing Models and Implications for SafeInsights
**Jiayi Zhang, Ryan S. Baker, Debshila Basu Mallick, Cristina Heffernan, Neil Heffernan** — cs.HC This paper replicates and extends prior work on the cold-start problem in knowledge tracing — the chal…
2026-06-10 · knowledge-tracing, learning-analytics, student-modeling, higher-ed, llm
📄 TurtleAI: Benchmarking Multimodal Models for Visual Programming in Turtle Graphics
> **Synthesis:** Vision-language models (VLMs) have been explored for visual programming, where they generate code to solve visual tasks. However, most prior work focuses on visual programming for pro…
2026-06-03 · cs-education, k-12, multimodal
📄 Benchmarking Large Language Models for Diagnosing Students' Cognitive Skills from Handwritten Math Work
> **MathCog** benchmark (3,036 teacher-annotated diagnostic verdicts, 639 handwritten responses, 18 LLMs): all models severely underperform (macro F1 < 0.5) — over-attributing evidence, overthinking m…
2026-05-31 · ai-ed-evaluation, knowledge-tracing, multimodal, human-in-the-loop, critical-thinking
🏷️ AI Ed Evaluation
> **AI-ed evaluation** — the body of methods, benchmarks, and criteria used to assess whether AI education tools (LLM-based tutors, automated graders, feedback systems, agents) actually work — not jus…
2026-05-29 · llm, assessment, formative-assessment, teacher-role, generative-ai
📄 How Students (Mis)understand Conditionals and Loops -- A Taxonomy
This paper presents a fine-grained taxonomy categorizing novice programmers' difficulties with reading and understanding control flow constructs — specifically conditionals (selection) and loops (iter…
2026-05-27 · cs-education, stem-education, student-experience, higher-ed, llm
📄 StanBKT: Rethinking Parameter Estimation in Bayesian Knowledge Tracing
StanBKT introduces an open-source Python package for Bayesian Knowledge Tracing (BKT) that moves beyond traditional expectation-maximization (EM) point estimates to full Bayesian inference via Stan. T…
2026-05-25 · intelligent-tutoring, learning-analytics, adaptive-learning, open-source, adaptive-learning-systems
📄 From Heuristics to Analytics: Forecasting Effort and Progress in Online Learning
This paper tackles a core ITS challenge: predicting when students will disengage so tutors can intervene before it's too late. It introduces **engagement forecasting** as a supervised prediction task …
2026-05-20 · intelligent-tutoring, learning-analytics, engagement-metrics, k-12, efficacy-study
📄 What Makes Words Hard? Sakura at BEA 2026 Shared Task on Vocabulary Difficulty Prediction
🔗 [Code](https://github.com/adno/vocabulary-difficulty) This paper presents two complementary approaches to predicting vocabulary difficulty for language learners, achieving state-of-the-art results …
2026-05-20 · language-learning, llm, generative-ai, scaffolding, formative-assessment
📄 Confirming Correct, Missing the Rest: LLM Tutoring Agents Struggle Where Feedback Matters Most
LLM tutors achieve near-ceiling on correct steps but systematically over-reject valid-suboptimal reasoning and over-validate incorrect solutions — precisely where adaptive tutoring matters most. This …
2026-05-19 · intelligent-tutoring, llm, generative-ai, scaffolding, feedback-loop
📄 Simulating Students or Sycophantic Problem Solving? On Misconception Faithfulness of LLM Simulators
This paper exposes a critical failure mode in using LLMs as simulated students for [[intelligent-tutoring]] development and evaluation. The authors introduce **misconception faithfulness** — the prope…
2026-05-16 · intelligent-tutoring, llm, generative-ai, hallucination-risk, student-experience
📄 Are Agents Ready to Teach? A Multi-Stage Benchmark for Real-World Teaching Workflows
> Are Agents Ready to Teach? A Multi-Stage Benchmark for Real-World Teaching Workflows **Chen et al. (2026)** — Multiple institutions. Under review.…
2026-05-15 · agentic-ai, generative-ai, intelligent-tutoring, llm, scaffolding
📄 Sycophancy is an Educational Safety Risk: Why LLM Tutors Need Sycophancy Benchmarks
> Sycophancy is an Educational Safety Risk: Why LLM Tutors Need Sycophancy Benchmarks **Kasneci & Kasneci (2026)** — Position paper. arXiv cs.AI/cs.HC.…
2026-05-15 · intelligent-tutoring, hallucination-risk, llm, generative-ai, over-reliance
📄 Ensuring Reliability in Programming Knowledge Tracing: A Re-evaluation of Attention-augmented Models and Experimental Protocols
This ITS 2026 paper challenges claims about attention-augmented Programming Knowledge Tracing (PKT) superiority. The authors identify three critical protocol flaws: **attention dimension misconfigurat…
2026-05-13 · knowledge-tracing, automated-grading, learning-analytics
📄 Reinforcement Learning Measurement Model
Interactive assessments generate sequential process data that conventional item response models (IRT) cannot adequately handle. This paper proposes a **reinforcement learning measurement model** that …
2026-05-12 · assessment, learning-analytics, knowledge-tracing, llm
📄 AcademiClaw: When Students Set Challenges for AI Agents
> **Yu, Lu, Si et al. (77 authors, 2026)** — Shanghai Jiao Tong University, SII, GAIR. Open-source benchmark.…
2026-05-11 · higher-ed, llm, generative-ai, student-experience, pedagogical-llm-training
📄 When AI Agents Teach Each Other: Discourse Patterns Resembling Peer Learning in the Moltbook Community
> **Authors:** Eason Chen, Ce Guan, A Elshafiey, Zhonghao Zhao, Joshua Zekeri, Afeez Edeifo Shaibu, Emmanuel Osadebe Prince **Year:** 2026 **Venue:** arXiv (cs.HC) > Mining discourse from Moltbook, a …
2026-05-11 · agentic-ai, collaborative-ai-tutoring, engagement-metrics, learning-analytics, llm
📄 The Missing Evaluation Axis: What 10,000 Student Submissions Reveal About AI Tutor Effectiveness
> A framework for evaluating AI tutoring systems that extends beyond pedagogical quality of feedback to measure what students actually *do* with that feedback — whether they act on it and whether they…
2026-05-09 · intelligent-tutoring, efficacy-study, higher-ed, engagement-metrics, llm
📄 NSMQ Riddles: A Benchmark of Scientific and Mathematical Riddles for Quizzing Large Language Models
> Boateng et al. (2026) introduce **NSMQ Riddles**, a benchmark of 1.8K scientific and mathematical riddles drawn from 11 years of Ghana's **National Science and Maths Quiz** — a live TV competition f…
2026-05-08 · stem-education, k-12, llm, efficacy-study, pedagogical-llm-training
📄 Quality-Conditioned Agreement in Automated Short Answer Scoring: Mid-Range Degradation and the Impact of Task-Specific Adaptation
> Schleifer, Ariely & Klebanov (2026) investigate a critical gap in [[automated-grading]]: **how scoring quality degrades for mid-range student responses**. Most ASAS evaluations focus on clearly corr…
2026-05-08 · automated-grading, formative-assessment, llm, efficacy-study, human-in-the-loop-ai
📄 TeachBench - Evaluating LLM Teaching Ability
> While LLMs are increasingly used as teaching assistants, their teaching capability remains insufficiently evaluated — a critical gap in current AIED research. > Syllabus-grounded framework for measu…
2026-05-08 · llm, formative-assessment, personalized-learning, feedback-loop, ai-literacy
📄 Agentic Workflows in Education
> A design framework for educational AI systems structured around four agentic paradigms: **reflection**, **planning**, **tool use**, and **multi-agent collaboration**. Proposed by Kamalov et al. (202…
2026-05-07 · agentic-ai, ai-education, intelligent-tutoring, pedagogical-llm-training, human-in-the-loop-ai
📄 AI Tutor Effectiveness Review
> Zerkouk, Mihoubi & Chikhaoui (2025) systematically analyzed qualified studies from 2010–2025 across: > A comprehensive systematic review of AI-based Intelligent Tutoring Systems (2010–2025) reveals …
2026-05-07 · intelligent-tutoring, efficacy-study, higher-ed, k-12, pedagogical-llm-training
📄 Educational LLM Alignment
> Hardy & Kim (2026) identify a **cascading proxy** problem in AI-for-education evaluation: > The gap between what LLMs are *capable* of and what actually *benefits learners* — benchmark performance, …
2026-05-07 · llm, efficacy-study, bias-mitigation, teacher-role, pedagogical-llm-training
📄 Educational VLM Evaluation
> Benchmarking vision-language models (VLMs) not on their ability to solve problems, but on their ability to *support learners* — particularly struggling learners and those making errors. Traditional …
2026-05-07 · assessment, multimodal, pedagogical-safety, stem-education, ai-education
📄 LLM-Based Educational Simulation: Evaluating Temporal Student Persona Stability Across ADHD Profiles
> Gonnermann-Müller, Haase & Leins (2026) evaluate whether **LLM-generated student personas simulating ADHD profiles** maintain stable and realistic behavioral patterns over time. This addresses a cri…
2026-05-07 · llm, student-experience, ai-education, generative-ai, learning-analytics
🏷️ Human-in-the-Loop AI for Education
Educational AI systems that strategically interleave automated generation with human expert judgment, preserving pedagogical quality while scaling production. Two recent implementations illustrate dis…
2026-05-07 · human-in-the-loop, assessment, pedagogical-safety, ai-education, llm
🏷️ Training Pedagogical LLMs for Tutoring
> Domain-specialized optimization can transform a mid-sized open-source model (Qwen3-32B) into a pedagogical domain expert that outperforms far larger proprietary systems — but only when training rewa…
2026-05-07 · llm, intelligent-tutoring, adaptive-learning, ai-education, higher-ed

Related Tags

llm (44)generative-ai (22)intelligent-tutoring (20)higher-ed (14)k-12 (14)scaffolding (12)formative-assessment (11)learning-analytics (11)agentic-ai (10)efficacy-study (10)rag (10)knowledge-tracing (10)student-experience (10)ai-education (9)automated-grading (9)