Are Agents Ready to Teach? A Multi-Stage Benchmark for Real-World Teaching Workflows Chen et al. (2026) โ Multiple institutions. Under review. ๐ Full text (arXiv)
Summary
EduAgentBench introduces the first theory-grounded, holistic benchmark for evaluating AI tutor agents across the full scope of real teaching work. Unlike existing benchmarks that focus narrowly on answer correctness, EduAgentBench defines 150 source-grounded tasks spanning three capability surfaces:
1. Professional pedagogical judgment โ making evidence-based instructional decisions aligned with intelligent-tutoring principles. 2. Situated multi-turn tutoring โ diagnosing learner state and adapting scaffolding over extended dialogue interactions. 3. Canvas-style teaching workflow completion โ executing tasks within realistic learning management systems (posting assignments, grading, providing feedback-loop).
The benchmark is constructed through a pedagogical-insight-driven pipeline with complementary human review and automatic verification signals. Evaluating frontier models reveals a critical gap: current LLMs demonstrate bounded pedagogical judgment but fall short of professional teaching standards in both situated tutoring and autonomous workflow execution. This connects directly to concerns about agentic-workflows-education and whether conversational-ai-tutors-framework can truly meet classroom demands.
The finding that models struggle most with multi-step teaching workflows in realistic environments echoes broader multi-agent-instructional-design challenges and the human-in-the-loop-ai requirements for production educational systems. The benchmark provides a measurement foundation for developing tutor agents that can genuinely support real teaching work, complementing existing evaluations like teachbench-llm-teaching-evaluation.
Related Pages
- codify-socratic-tutoring-programming โ Socratic tutoring platform for programming with integrated assessment
- ai-tpack-teacher-multi-agent-workflow โ How teachers design and orchestrate multi-agent instructional workflows
- intelligent-tutoring โ Core AI tutoring systems and architectures
- teachbench-llm-teaching-evaluation โ Complementary LLM teaching ability benchmark
- agentic-workflows-education โ Agentic workflows in educational contexts
- conversational-ai-tutors-framework โ Framework for conversational AI tutoring
- pedagogical-llm-training โ Training LLMs for pedagogical competence
- scaffolding โ Adaptive support in learning
- ai-tutor-behavioral-evaluation โ Behavioral evaluation of AI tutors from student data
- retrieval-augmented-tutoring-algorithm-kite โ KITE: complementary tutoring architecture with simulated evaluation