🧠 AI Ed Wiki

🏷️ llm-evaluation

5 pages tagged with llm-evaluation(5 articles, 0 concepts)

📄 ELBench: A Multi-Dimensional Benchmark for Education-Facing Large Language Models
> **Synthesis:** Jiang et al. (2026) introduce **ELBench**, the first benchmark to evaluate education-facing LLMs on all four required dimensions — General Capability, Safety and Trustworthiness, Basi…
2026-08-13 · benchmark, ai-ed-evaluation, pedagogical-safety, llm, safety
📄 Findings of the First Teaching Monster Challenge: A Benchmark of Pedagogical Content Knowledge in AI Agents
> **Synthesis:** Lin et al. (2026) present the **Teaching Monster Challenge**, the first instructional-video generation benchmark that treats the learner persona as an explicit evaluation criterion, m…
2026-08-13 · benchmark, ai-ed-evaluation, agentic-ai, pedagogical-agent, content-generation
📄 From Evaluated Models to Evaluation Aids: A Multi-Evidence Study of LLM-Based Difficulty Calibration for Programming Examinations
> **Synthesis:** Yan, Xiong, Li & Chen (2026) reposition LLMs from benchmark targets to auxiliary evidence sources for interpreting programming-exam difficulty, showing that AI difficulty estimates co…
2026-08-11 · computing-education, programming-education, assessment, automated-assessment, educational-measurement
📄 When Help is Unhelpful: Evaluating AI Tutors for Productive Struggle
> **Synthesis:** Zhang et al. (2026) introduce TutorMoments, a replay-based evaluation framework that tests whether LM tutors adapt their pedagogical actions to context — scaffolding when support is n…
2026-08-08 · intelligent-tutoring, scaffolding, llm, k-12, math-education
📄 Can LLMs Effectively Simulate Human Learners? Teachers' Insights from Tutoring LLM Students
> **Synthesis:** Semi-structured interviews with 12 teachers who tutored LLM-simulated students (MathDial dataset) reveal key authenticity gaps: overly complex language, lack of emotions, unnatural at…
2026-08-06 · llm, student-simulation, teacher-training, dialogue-tutoring, k-12

Related Tags

benchmark (3)llm (3)ai-ed-evaluation (2)generative-ai (2)scaffolding (2)k-12 (2)pedagogical-safety (1)safety (1)agentic-ai (1)pedagogical-agent (1)content-generation (1)computing-education (1)programming-education (1)assessment (1)automated-assessment (1)