🧠 AI Ed Wiki

Are Agents Ready to Teach? A Multi-Stage Benchmark for Real-World Teaching Workflows Chen et al. (2026) — Multiple institutions. Under review.

Are Agents Ready to Teach? A Multi-Stage Benchmark for Real-World Teaching Workflows

Summary

EduAgentBench introduces the first theory-grounded, holistic benchmark for evaluating AI tutor agents across the full scope of real teaching work. Unlike existing benchmarks that focus narrowly on answer correctness, EduAgentBench defines 150 source-grounded tasks spanning three capability surfaces:

1. Professional pedagogical judgment — making evidence-based instructional decisions aligned with Intelligent Tutoring principles.

2. Situated multi-turn tutoring — diagnosing learner state and adapting Scaffolding over extended dialogue interactions.

3. Canvas-style teaching workflow completion — executing tasks within realistic learning management systems (posting assignments, grading, providing Feedback Loop).

The benchmark is constructed through a pedagogical-insight-driven pipeline with complementary human review and automatic verification signals. Evaluating frontier models reveals a critical gap: current LLMs demonstrate bounded pedagogical judgment but fall short of professional teaching standards in both situated tutoring and autonomous workflow execution. This connects directly to concerns about Agentic Workflows Education and whether Conversational AI Tutors Framework can truly meet classroom demands.

The finding that models struggle most with multi-step teaching workflows in realistic environments echoes broader Multi Agent Instructional Design challenges and the Human In The Loop AI requirements for production educational systems. The benchmark provides a measurement foundation for developing tutor agents that can genuinely support real teaching work, complementing existing evaluations like Teachbench LLM Teaching Evaluation.

Connected Concepts

  • Intelligent Tutoring
  • Scaffolding
  • Feedback Loop
  • Human In The Loop AI
  • Connected Articles

  • Agentic Workflows Education
  • Conversational AI Tutors Framework
  • Multi Agent Instructional Design
  • Teachbench LLM Teaching Evaluation
  • Citation

    Chen, Z., Liu, P., Sheng, R., Li, H., Tu, J., Deng, X., Shum, K., Liu, D., & Qu, H. (2026). Are agents ready to teach? A multi-stage benchmark for real-world teaching workflows. arXiv:2605.14322.