Evaluation benchmarks for AI systems in educational contexts
Connections
Related Pages
- llm-cognitive-diagnosis-handwritten-math โ MathCog benchmark: 18 LLMs evaluated on cognitive skill diagnosis from handwritten math; all F1 < 0.5; systematic over-attribution and hallucination of evidence (2025)
- vocabulary-difficulty-prediction โ LLM fine-tuned with soft-target loss achieves r>0.91 for vocabulary difficulty p
- reliable-programming-kt โ Standardized evaluation protocol for programming knowledge tracing
- nsmq-riddles-science-math-benchmark โ First educational AI benchmark originating from the Global South
- short-answer-scoring-quality-degradation โ Quality-conditioned evaluation as a new benchmark dimension for ASAS
- llm-student-simulation-misconception-faithfulness โ Selective Flip Score as new metric for student simulator evaluation- measuring-llm-tutors-teach-vs-solve -- Benchmarks conflate solving with teaching (r=0.421); pedagogy and solving scores should be reported separately
- rethinking-scaffolding-llm-tutors โ Rethinking Scaffolding in LLM Tutors
- mllm-scientific-visualization-literacy โ reusable SciVis-literacy benchmark methodology (49 items, 18 viz)
- llm-psychometric-calibration-cdp โ CDP framework dramatically improves LLM-simulated examinee alignment with human ...
๐ 43 other pages tagged benchmark
- AcademiClaw: When Students Set Challenges for AI Agents
- Agentic Workflows in Education
- AI Tutor Effectiveness Review
- Aligning LLM-Simulated and Human Examinees for Psychometric Calibration: A Cognitive Diagnostic Profiling Approach
- Are Agents Ready to Teach? A Multi-Stage Benchmark for Real-World Teaching Workflows
- Benchmarking Large Language Models for Diagnosing Students' Cognitive Skills from Handwritten Math Work
- Benchmarking Multimodal Large Language Models for Scientific Visualization Literacy
- Cognitive Episodes in LLM Reasoning Traces Enable Interpretable Human Item Difficulty Prediction
- Confirming Correct, Missing the Rest: LLM Tutoring Agents Struggle Where Feedback Matters Most
- CoTAL: Human-in-the-Loop Prompt Engineering for Generalizable Formative Assessment Scoring and Feedback
- CSTutorBench: Benchmarking Small Language Models as Tutors for Block-Based Programming
- Educational LLM Alignment
- Educational VLM Evaluation
- EduGuard: A Safe RAG-Based LLM Tutor for Programming Education
- EduZone: A Framework for Evaluating LLM Safety for K-12 Students and Teachers
- Ensuring Reliability in Programming Knowledge Tracing: A Re-evaluation of Attention-augmented Models and Experimental Protocols
- From Heuristics to Analytics: Forecasting Effort and Progress in Online Learning
- How Students (Mis)understand Conditionals and Loops -- A Taxonomy
- ICLE++: Modeling Fine-Grained Traits for Holistic Essay Scoring
- ISD Agent Benchmark
- Learning Analytics
- Learning Engagement Assistant (LEA): Cross-Course Scalability and Classroom Evaluation of an Agentic AI Tutoring System
- LLM-Based Educational Simulation: Evaluating Temporal Student Persona Stability Across ADHD Profiles
- Measuring Whether LLM Tutors Teach or Solve: A Diagnostic for Educational Impact
- MedGame: Storytelling Gamification Empowered by Large Language Models for Medical Education
- Modernizing Ground Truth: Four Shifts Toward Improving Reliability and Validity in AI in Education
- NSMQ Riddles: A Benchmark of Scientific and Mathematical Riddles for Quizzing Large Language Models
- Quality-Conditioned Agreement in Automated Short Answer Scoring: Mid-Range Degradation and the Impact of Task-Specific Adaptation
- Reexamining the Cold-Start Problem in Knowledge Tracing Models and Implications for SafeInsights
- Reinforcement Learning Measurement Model
- Representation Robustness under Executable Reasoning Constraints in Large Language Models for Mathematical Problem Solving
- Rethinking LLM-Judged Helpfulness as a Pedagogy Signal: A Pre-Registered Audit Across Tutor Models
- Rethinking Scaffolding in LLM Tutors: The Interactional Mismatch Between Benchmarks and Real-World Deployments
- Simulating Students or Sycophantic Problem Solving? On Misconception Faithfulness of LLM Simulators
- StanBKT: Rethinking Parameter Estimation in Bayesian Knowledge Tracing
- Sycophancy is an Educational Safety Risk: Why LLM Tutors Need Sycophancy Benchmarks
- TeachBench - Evaluating LLM Teaching Ability
- The Missing Evaluation Axis: What 10,000 Student Submissions Reveal About AI Tutor Effectiveness
- The Tutoring Effectiveness Index: Predicting LLM Math Tutor Quality from Four Conversation Signals
- Training Pedagogical LLMs for Tutoring
- TurtleAI: Benchmarking Multimodal Models for Visual Programming in Turtle Graphics
- What Makes Words Hard? Sakura at BEA 2026 Shared Task on Vocabulary Difficulty Prediction
- When AI Agents Teach Each Other: Discourse Patterns Resembling Peer Learning in the Moltbook Community