Research Article
TeachBench - Evaluating LLM Teaching Ability
Synthesis: While LLMs are increasingly used as teaching assistants, their teaching capability remains insufficiently evaluated — a critical gap in current AIED research.
Syllabus-grounded framework for measuring Large Language Models (LLMs) teaching capability via student performance improvement after multi-turn instruction.
The Gap in LLM Evaluation
Li et al. (2026) address a critical gap: while LLMs are increasingly used as teaching assistants, their teaching capability remains insufficiently evaluated.
Limitations of Existing Benchmarks
| Benchmark Type | Focus | Limitation |
|---|---|---|
| Problem Solving (MMLU, HELM, GSM8K) | Answer correctness | Measures solver, not teacher |
| Exam-centric (AGIEval, C-Eval, GAOKAO-Bench) | Exam performance | Still solution-centric |
| Tutoring dialogues (MathDial, TutorBench) | Per-turn response quality | Misses end-to-end teaching effectiveness |
TeachBench shifts the evaluation target: from solving to teaching.
Syllabus-Grounded Framework
Core Design Principles
- Knowledge-centered: Evaluation based on structured knowledge points (syllabus), not target questions
- Leakage-controlled: Teacher agents restricted to knowledge points + example problems (no access to test items)
- Outcome-based: Teaching effectiveness measured by student agent's performance improvement
Workflow
Syllabus → Knowledge Tree → Teacher Agent (multi-turn instruction) → Student Agent → Performance Gain
Key Findings from Gaokao Experiments
Using Chinese National College Entrance Examination (Gaokao) data across multiple subjects:
| Finding | Implication |
|---|---|
| Domain variation: Math teaching effective (7.63pt gain with Qwen3-235B), but physics/chemistry challenging | Teaching ability is domain-specific, not generalized |
| Example problems backfire: Models shift to error correction vs. syllabus-grounded instruction | Current LLMs struggle with structured teaching vs. reactive problem-solving |
| Teaching ≠ Solving: Models good at solving aren't necessarily good at teaching | Teaching ability is a distinct LLM behavior dimension |
Connection to Existing Work
vs. AI Tutor Effectiveness
- Traditional ITS effectiveness reviews focus on human learning outcomes with deployed systems
- TeachBench evaluates model teaching capability in controlled agentic settings
- Both highlight: teaching is more than problem-solving
vs. Educational LLM Alignment
- Alignment benchmarks measure: "Does this model produce good teaching content?"
- TeachBench measures: "Does this model improve learning through instruction?"
- Complementary: alignment → content quality; TeachBench → instructional effectiveness
vs. Agentic Workflows
- TeachBench operationalizes the "teacher agent" paradigm in agentic education
- Reveals current LLMs struggle with structured pedagogical planning (vs. reactive Q&A)
- Aligns with: agentic reflection, planning, and tool use in educational contexts
What this means for practice
- Researchers. Evaluate teaching as a distinct capability rather than inferring it from solving: syllabus-grounded evaluation measures whether a model improves a learner's performance, whereas Problem Solving and exam-centric benchmarks such as MMLU, GSM8K, and AGIEval score answer correctness alone.
- Researchers. Restrict teacher agents to syllabus knowledge points and example problems and hold the student agent fixed, so the model cannot be handed the items it is meant to teach and results stay reproducible across runs.
- Instructors. Do not carry a model's teaching strength across subjects: the strongest result was a 7.63-point gain in mathematics with Qwen3-235B-A22B-Instruct, while physics and chemistry showed the weakest teaching outcomes.
- Designers. Test whether worked examples help before attaching them to a tutor: incorporating example problems shifted models toward example-specific error correction instead of syllabus-grounded instruction.
- Instructors. Judge a tutoring system on learning gains accumulated over multi-turn instruction, not on per-turn response quality or user satisfaction.
Limitations
- Teaching is measured with LLM-based student agents standing in as proxies for human learners; the authors state this controlled setting may not fully reflect the diversity and complexity of human learning behaviors.
- No human teachers were included as a baseline, so the experiments rank models against one another rather than against human instructional performance.
- The study of example-based teaching is limited to a specific interaction design; the authors note that alternative instructional protocols may produce different outcomes.
- The benchmark is built from Gaokao (Chinese National College Entrance Examination) syllabi and questions across seven subjects — Mathematics, Physics, Chemistry, Biology, History, Geography, and Politics — so the domain rankings are tied to that exam's knowledge structure.
Citation
Li, Z., Song, S., Ma, J., Li, R., Zeng, Y., Li, M., et al. (2026). TeachBench - Evaluating LLM Teaching Ability.