📄 Research Article
ELBench: A Multi-Dimensional Benchmark for Education-Facing Large Language Models
Synthesis: Jiang et al. (2026) introduce ELBench, the first benchmark to evaluate education-facing LLMs on all four required dimensions — General Capability, Safety and Trustworthiness, Basic Education, and High-Level Cultivation — under a common protocol, combining curated public sources with newly synthesized safety and cultivation data. Testing nine models, they find module-level profiles are more informative than a single aggregate: the top six models are statistically indistinguishable overall yet differ substantially by module leader, and safety is anti-correlated with practical teaching (r = −0.83). The two education-specialized models lead neither education module, and all models share a systematic blind spot on High-Level Cultivation's structured-judgment task. The work connects to Benchmark, AI Ed Evaluation, and AI Ed Evaluation frameworks.
An Integrated Profile, Not a Single Score
A usable education-facing model must be accurate, safe under sensitive prompts, instructionally useful, and aligned with pedagogical goals at the same time. Existing benchmarks evaluate these requirements largely in isolation; ELBench is the first to assess all four as an integrated profile under a common protocol.
Three Findings
Connected Concepts
Connected Articles
Citation
Jiang, Y., Zhu, X., Tan, F., Zhang, Z., Huang, K., Yu, Y., Fei, Z., Luo, Y., Li, K., Hao, H., Zhai, G., & Zhou, A. (2026). ELBench: A multi-dimensional benchmark for education-facing large language models. arXiv:2608.09548.