EduZone is an automated evaluation framework that generates contextually grounded adversarial interactions to probe LLM safety in K-12 education, revealing that models are more vulnerable to education-specific harms and dynamic multi-turn conversations than existing guardrails address.
Junyeong Park, Jieun Han, Haneul Yoo, So-Yeon Ahn, Jinsung Yoon, Alice Oh โ arXiv (cs.CY / cs.AI) preprint, 2026 (KAIST, Google Cloud AI Research, NYU). ๐ Full text (arXiv)
Synthesis
Combines student- and teacher-facing LLM usage contexts with fine-grained curriculum concepts and 6 risk categories / 28 subcategories spanning conventional and education-specific harms.
Builds adversarial interactions in three settings: single-turn requests, static multi-turn conversations, and dynamic multi-turn conversations.
Evaluates ten LLMs across four safety levels: refusal, safe assistance, risky assistance with safety guidance, and fully risky assistance.
Results show greater vulnerability to education-specific risks and dynamic multi-turn interactions; existing safety guardrails fail to adequately address these risks.
Related Pages
- ai-tutor-safety-harms โ safety risks specific to AI tutoring systems
- pedagogical-safety โ safety of AI use within pedagogical contexts
- k-12 โ K-12 education as the deployment context
- ai-governance-education โ governance frameworks for AI in education
- ai-k12-evidence-base โ the evidence base for AI in K-12
- llm โ large language models as the evaluated systems
Citation
APA: Junyeong Park, Jieun Han, Haneul Yoo, So-Yeon Ahn, Jinsung Yoon, Alice Oh (2026). EduZone: A Framework for Evaluating LLM Safety for K-12 Students and Teachers. arXiv:2608.02024. arXiv (cs.CY / cs.AI) preprint.