🏷️ benchmark
54 pages tagged with benchmark(49 articles, 5 concepts)
📄 ELBench: A Multi-Dimensional Benchmark for Education-Facing Large Language Models
> **Synthesis:** Jiang et al. (2026) introduce **ELBench**, the first benchmark to evaluate education-facing LLMs on all four required dimensions — General Capability, Safety and Trustworthiness, Basi…
📄 Findings of the First Teaching Monster Challenge: A Benchmark of Pedagogical Content Knowledge in AI Agents
> **Synthesis:** Lin et al. (2026) present the **Teaching Monster Challenge**, the first instructional-video generation benchmark that treats the learner persona as an explicit evaluation criterion, m…
🏷️ Research Methods in AIED
> **Research methods in AIED** — the set of empirical designs, data-collection strategies, and analytic techniques researchers use to study AI in education: whether and how AI tools support (or harm) …
📄 Adversarial Stress Testing of Role-Playing Language Agents using Multi-Agent Evaluation
> **Synthesis:** This paper presents a modular multi-agent platform for adversarially stress-testing [[agentic-ai|role-playing language agents]] through structured multi-turn dialogue. With three coor…
🏷️ Benchmark
> **Benchmark** — standardized test suites and evaluation frameworks used to measure AI model performance on educational tasks. Benchmarks enable reproducible comparison across models and approaches, …
📄 When Help is Unhelpful: Evaluating AI Tutors for Productive Struggle
> **Synthesis:** Zhang et al. (2026) introduce TutorMoments, a replay-based evaluation framework that tests whether LM tutors adapt their pedagogical actions to context — scaffolding when support is n…
📄 What out-of-the-box LLMs can(t) do in law? A Turing test in Italian exams for lawyers, judges and notaries
> **Synthesis:** This paper reports a blind Turing Test evaluating leading LLMs on three Italian professional legal examinations: the Bar exam, Judges exam, and Notary exam. LLMs generated full writte…
📄 Robust and Efficient Motion Reasoning for Privacy-Aware Classroom Incident Recognition
> **Synthesis:** Pilot study on privacy-aware computer vision for classroom incident detection. Introduces a hybrid benchmark combining generative CCTV-style videos with real classroom pose data. Prop…
📄 When Agents Learn to Be You: Benchmarking Privacy Leakage, Impersonation Risk, and Defenses in Persona Skills
> **When Agents Learn to Be You: Benchmarking Privacy Leakage, Impersonation Risk, and Defenses in Persona Skills** — Introduces AntiSkillBench with 7,500 persona-grounded dialogue traces from 50 beha…
📄 EduClaw-Bench: A Long-Horizon Benchmark for Pedagogical LLM Agents with Simulated Learners
> **EduClaw-Bench: A Long-Horizon Benchmark for Pedagogical LLM Agents with Simulated Learners** — Introduces a 30-day long-horizon benchmark for pedagogical LLM agents using simulated learners ground…
📄 EduZone: A Framework for Evaluating LLM Safety for K-12 Students and Teachers
> **EduZone is an automated evaluation framework that generates contextually grounded adversarial interactions to probe LLM safety in K-12 education, revealing that models are more vulnerable to educa…
📄 CoTAL: Human-in-the-Loop Prompt Engineering for Generalizable Formative Assessment Scoring and Feedback
> 1. **Evidence-Centered Design (ECD)** — assessments and rubrics aligned to curriculum goals from the start 2. **Human-in-the-loop prompt engineering** — labelled examples and prompts refined iterati…
📄 ICLE++: Modeling Fine-Grained Traits for Holistic Essay Scoring
Introduces ICLE++, a new annotated corpus of persuasive student essays that addresses critical limitations of the dominant ASAP benchmark in [[automated-essay-scoring]] research. Unlike ASAP — used by…
📄 Rethinking LLM-Judged Helpfulness as a Pedagogy Signal: A Pre-Registered Audit Across Tutor Models
Pre-registered study auditing whether general-purpose helpfulness rubrics can distinguish direct answer-giving from pedagogical guidance in LLM tutors. Uses deterministic detectors for answer leakage …
📄 Aligning LLM-Simulated and Human Examinees for Psychometric Calibration: A Cognitive Diagnostic Profiling Approach
Proposes Cognitive Diagnostic Profiling (CDP), a zero-shot framework that dramatically improves LLM-simulated examinee alignment with human test-takers. With CDP, IRT difficulty Spearman correlations …
📄 ISD Agent Benchmark
> **ISD-Agent-Bench** is a comprehensive benchmark for evaluating LLM-based instructional design agents, comprising **25,795 scenarios** generated via a Context Matrix framework that combines 51 conte…
📄 MedGame: Storytelling Gamification Empowered by Large Language Models for Medical Education
MedGame transforms static clinical cases into structured, executable storytelling games for medical education, moving beyond the localized question-answering and single-turn feedback that characterize…
📄 Representation Robustness under Executable Reasoning Constraints in Large Language Models for Mathematical Problem Solving
This study probes how sensitive [[llm]] mathematical problem solving is to the surface representation of an item — a question with direct bearing on [[assessment-validity]] when LLMs are used for scor…
📄 EduGuard: A Safe RAG-Based LLM Tutor for Programming Education
EduGuard is a retrieval-augmented generation (RAG) tutoring framework that directly confronts the safety and pedagogical failures of unrestricted LLM tutors in introductory programming. Unrestricted t…
📄 Benchmarking Multimodal Large Language Models for Scientific Visualization Literacy
Multimodal large language models (MLLMs) are increasingly used to interpret visualizations, yet most evaluations remain chart-centric and offer limited insight into **scientific visualization (SciVis)…
📄 Learning Engagement Assistant (LEA): Cross-Course Scalability and Classroom Evaluation of an Agentic AI Tutoring System
LEA (Learning Engagement Assistant) is an **agentic AI tutoring system** that couples course-specific retrieval-augmented generation (RAG) with structured [[knowledge-tracing]] / Knowledge Component (…
📄 CSTutorBench: Benchmarking Small Language Models as Tutors for Block-Based Programming
Deploying LLM tutors in K-12 raises concerns around privacy, cost, and reliance on proprietary models, motivating small language models (SLMs) as an alternative. The authors introduce **CSTutorBench**…
📄 Cognitive Episodes in LLM Reasoning Traces Enable Interpretable Human Item Difficulty Prediction
This paper introduces Epi2Diff (Episode to Difficulty), a framework that maps LLM reasoning traces into cognitively grounded episode sequences for predicting human item difficulty in [[assessment|educ…
📄 Rethinking Scaffolding in LLM Tutors: The Interactional Mismatch Between Benchmarks and Real-World Deployments
> **Alexandra Neagu, Jeffrey T. H. Wong, Marcus Messer, Rhodri Nelson, Peter B. Johnson** (2026). Pluralistic Alignment Workshop @ ICML 2026…
📄 Measuring Whether LLM Tutors Teach or Solve: A Diagnostic for Educational Impact
Studies whether public LLM tutoring benchmarks distinguish learning-supportive behavior from mere answer production. Proposes a lightweight diagnostic based on the gap between solving-oriented and ped…
📄 Reexamining the Cold-Start Problem in Knowledge Tracing Models and Implications for SafeInsights
**Jiayi Zhang, Ryan S. Baker, Debshila Basu Mallick, Cristina Heffernan, Neil Heffernan** — cs.HC This paper replicates and extends prior work on the cold-start problem in knowledge tracing — the chal…
📄 TurtleAI: Benchmarking Multimodal Models for Visual Programming in Turtle Graphics
> **Synthesis:** Vision-language models (VLMs) have been explored for visual programming, where they generate code to solve visual tasks. However, most prior work focuses on visual programming for pro…
📄 The Tutoring Effectiveness Index: Predicting LLM Math Tutor Quality from Four Conversation Signals
> **Authors:** Shim Jaechang, Unggi Lee (2026) — CIKM 2026…
📄 Benchmarking Large Language Models for Diagnosing Students' Cognitive Skills from Handwritten Math Work
> **MathCog** benchmark (3,036 teacher-annotated diagnostic verdicts, 639 handwritten responses, 18 LLMs): all models severely underperform (macro F1 < 0.5) — over-attributing evidence, overthinking m…
🏷️ AI Ed Evaluation
> **AI-ed evaluation** — the body of methods, benchmarks, and criteria used to assess whether AI education tools (LLM-based tutors, automated graders, feedback systems, agents) actually work — not jus…
📄 How Students (Mis)understand Conditionals and Loops -- A Taxonomy
This paper presents a fine-grained taxonomy categorizing novice programmers' difficulties with reading and understanding control flow constructs — specifically conditionals (selection) and loops (iter…
📄 StanBKT: Rethinking Parameter Estimation in Bayesian Knowledge Tracing
StanBKT introduces an open-source Python package for Bayesian Knowledge Tracing (BKT) that moves beyond traditional expectation-maximization (EM) point estimates to full Bayesian inference via Stan. T…
📄 From Heuristics to Analytics: Forecasting Effort and Progress in Online Learning
This paper tackles a core ITS challenge: predicting when students will disengage so tutors can intervene before it's too late. It introduces **engagement forecasting** as a supervised prediction task …
📄 What Makes Words Hard? Sakura at BEA 2026 Shared Task on Vocabulary Difficulty Prediction
🔗 [Code](https://github.com/adno/vocabulary-difficulty) This paper presents two complementary approaches to predicting vocabulary difficulty for language learners, achieving state-of-the-art results …
📄 Confirming Correct, Missing the Rest: LLM Tutoring Agents Struggle Where Feedback Matters Most
LLM tutors achieve near-ceiling on correct steps but systematically over-reject valid-suboptimal reasoning and over-validate incorrect solutions — precisely where adaptive tutoring matters most. This …
📄 Simulating Students or Sycophantic Problem Solving? On Misconception Faithfulness of LLM Simulators
This paper exposes a critical failure mode in using LLMs as simulated students for [[intelligent-tutoring]] development and evaluation. The authors introduce **misconception faithfulness** — the prope…
📄 Are Agents Ready to Teach? A Multi-Stage Benchmark for Real-World Teaching Workflows
> Are Agents Ready to Teach? A Multi-Stage Benchmark for Real-World Teaching Workflows **Chen et al. (2026)** — Multiple institutions. Under review.…
📄 Sycophancy is an Educational Safety Risk: Why LLM Tutors Need Sycophancy Benchmarks
> Sycophancy is an Educational Safety Risk: Why LLM Tutors Need Sycophancy Benchmarks **Kasneci & Kasneci (2026)** — Position paper. arXiv cs.AI/cs.HC.…
📄 Ensuring Reliability in Programming Knowledge Tracing: A Re-evaluation of Attention-augmented Models and Experimental Protocols
This ITS 2026 paper challenges claims about attention-augmented Programming Knowledge Tracing (PKT) superiority. The authors identify three critical protocol flaws: **attention dimension misconfigurat…
📄 Reinforcement Learning Measurement Model
Interactive assessments generate sequential process data that conventional item response models (IRT) cannot adequately handle. This paper proposes a **reinforcement learning measurement model** that …
📄 AcademiClaw: When Students Set Challenges for AI Agents
> **Yu, Lu, Si et al. (77 authors, 2026)** — Shanghai Jiao Tong University, SII, GAIR. Open-source benchmark.…
📄 When AI Agents Teach Each Other: Discourse Patterns Resembling Peer Learning in the Moltbook Community
> **Authors:** Eason Chen, Ce Guan, A Elshafiey, Zhonghao Zhao, Joshua Zekeri, Afeez Edeifo Shaibu, Emmanuel Osadebe Prince **Year:** 2026 **Venue:** arXiv (cs.HC) > Mining discourse from Moltbook, a …
📄 Modernizing Ground Truth: Four Shifts Toward Improving Reliability and Validity in AI in Education
> Stop treating κ > 0.8 as a binary stamp of approval.…
📄 The Missing Evaluation Axis: What 10,000 Student Submissions Reveal About AI Tutor Effectiveness
> A framework for evaluating AI tutoring systems that extends beyond pedagogical quality of feedback to measure what students actually *do* with that feedback — whether they act on it and whether they…
📄 NSMQ Riddles: A Benchmark of Scientific and Mathematical Riddles for Quizzing Large Language Models
> Boateng et al. (2026) introduce **NSMQ Riddles**, a benchmark of 1.8K scientific and mathematical riddles drawn from 11 years of Ghana's **National Science and Maths Quiz** — a live TV competition f…
📄 Quality-Conditioned Agreement in Automated Short Answer Scoring: Mid-Range Degradation and the Impact of Task-Specific Adaptation
> Schleifer, Ariely & Klebanov (2026) investigate a critical gap in [[automated-grading]]: **how scoring quality degrades for mid-range student responses**. Most ASAS evaluations focus on clearly corr…
📄 TeachBench - Evaluating LLM Teaching Ability
> While LLMs are increasingly used as teaching assistants, their teaching capability remains insufficiently evaluated — a critical gap in current AIED research. > Syllabus-grounded framework for measu…
📄 Agentic Workflows in Education
> A design framework for educational AI systems structured around four agentic paradigms: **reflection**, **planning**, **tool use**, and **multi-agent collaboration**. Proposed by Kamalov et al. (202…
📄 AI Tutor Effectiveness Review
> Zerkouk, Mihoubi & Chikhaoui (2025) systematically analyzed qualified studies from 2010–2025 across: > A comprehensive systematic review of AI-based Intelligent Tutoring Systems (2010–2025) reveals …
📄 Educational LLM Alignment
> Hardy & Kim (2026) identify a **cascading proxy** problem in AI-for-education evaluation: > The gap between what LLMs are *capable* of and what actually *benefits learners* — benchmark performance, …
📄 Educational VLM Evaluation
> Benchmarking vision-language models (VLMs) not on their ability to solve problems, but on their ability to *support learners* — particularly struggling learners and those making errors. Traditional …
📄 LLM-Based Educational Simulation: Evaluating Temporal Student Persona Stability Across ADHD Profiles
> Gonnermann-Müller, Haase & Leins (2026) evaluate whether **LLM-generated student personas simulating ADHD profiles** maintain stable and realistic behavioral patterns over time. This addresses a cri…
🏷️ Human-in-the-Loop AI for Education
Educational AI systems that strategically interleave automated generation with human expert judgment, preserving pedagogical quality while scaling production. Two recent implementations illustrate dis…
🏷️ Training Pedagogical LLMs for Tutoring
> Domain-specialized optimization can transform a mid-sized open-source model (Qwen3-32B) into a pedagogical domain expert that outperforms far larger proprietary systems — but only when training rewa…