🧠 AI Ed Wiki

SafeTutors is a benchmark that jointly evaluates safety and pedagogy in AI tutoring systems across mathematics, physics, and chemistry. It argues that tutoring safety is fundamentally different from conventional LLM safety: the primary risk is not toxic content but the quiet erosion of learning through answer over-disclosure, misconception reinforcement, and the abdication of scaffolding.

Key Findings

1. Universal harm across all models. Every evaluated model β€” from 3.8B open-weight models to GPT-5-mini β€” exhibited broad pedagogical harm. No model was consistently safe across all subjects and interaction modes, indicating that tutoring safety is not solved by general capability improvements.

2. Scale does not reliably improve safety. Increasing model size did not produce consistent improvements in pedagogical safety, challenging the assumption that larger models are inherently better tutors. This finding parallels broader critiques in llm-evaluation that general benchmarks do not capture domain-specific safety requirements.

3. Multi-turn dialogue dramatically worsens behavior. Pedagogical failure rates escalate from 17.7% in single-turn interactions to 77.8% in multi-turn conversations. The crescendo-based escalation design reveals that models which appear safe in one-turn evaluations systematically degrade across sustained interaction β€” single-turn "safe/helpful" results mask systematic tutor failure.

4. Harms are subject-dependent. Violation patterns vary significantly across mathematics, physics, and chemistry, indicating that mitigations must be discipline-aware. A tutoring safety strategy that works for math may not transfer to science domains.

5. An 11-dimension, 48-sub-risk taxonomy grounds the evaluation. SafeTutors' risk taxonomy spans Cognitive, Epistemic, Metacognitive, Motivational-Affective, Developmental & Equity, Instructional Alignment, Behavioral & Inquiry, Ethical-Epistemic Integrity, Informational-Semantic, Reflective-Critical, and Pedagogical Relationship dimensions β€” each with multiple sub-risks drawn from learning-science literature.

Implications

SafeTutors fundamentally reframes the conversation around Pedagogical Safety and AI Tutor Safety Harms. The dominant paradigm has been to evaluate AI tutors on problem-solving accuracy and generic safety (toxicity, refusal), but SafeTutors demonstrates that a tutor can be technically accurate and "safe" by conventional metrics while systematically undermining learning. The benchmark's central insight β€” that tutoring harm is qualitatively different from content harm β€” has major implications for AI Tutoring regulation and deployment.

The multi-turn degradation finding is particularly alarming for real-world deployment. Most tutoring interactions extend over multiple turns, yet the evaluation community has largely relied on single-turn benchmarks. SafeTutors provides evidence that this practice is dangerously misleading. Systems like Eduzone LLM Safety K12 and Vetting Dual LLM Safety Education that prioritize multi-turn safety evaluation are essential, not optional.

The risk taxonomy itself is a significant contribution, providing a theoretically grounded vocabulary for discussing tutoring harm. It bridges educational-theory and AI safety, enabling researchers to move beyond vague claims about "tutor quality" toward precise identification of specific failure modes. This taxonomy could inform the design of Pedagogical Safety RL approaches like Singh Eduqwen Pedagogical RL 2026 that train models to avoid specific pedagogical harms.

For K 12 contexts, where the stakes of pedagogical harm are highest, SafeTutors provides empirical evidence that current models are not safe enough for unsupervised deployment. The subject-dependence of harms suggests that safety evaluation must be integrated into discipline-specific AI Tutor Behavioral Evaluation pipelines rather than treated as a one-time gate.

Connected Concepts

  • AI Tutoring
  • K 12
  • Pedagogical Safety
  • LLM
  • Regulation
  • Scaffolding
  • Connected Articles

  • AI Tutor Behavioral Evaluation β€” The Missing Evaluation Axis: What 10,000 Student Submissions Reveal About AI Tutor Effectiveness
  • AI Tutor Safety Harms β€” AI Tutor Safety and Pedagogical Harms
  • Eduzone LLM Safety K12 β€” EduZone: A Framework for Evaluating LLM Safety for K-12 Students and Teachers
  • Pedagogical Safety RL β€” Pedagogical Safety in Educational Reinforcement Learning
  • Singh Eduqwen Pedagogical RL 2026 β€” EduQwen: Pedagogical RL
  • Vetting Dual LLM Safety Education β€” VETTING: A dual-LLM framework for in-loop safety verification via policy isolation in educational AI
  • Aaai2026 Prompting Literacy K12 β€” Learning to Use AI for Learning: Teaching Responsible Use of AI Chatbot to K-12 Students Through an AI Literacy Module
  • Academiclaw Student Agent Benchmark β€” AcademiClaw: When Students Set Challenges for AI Agents
  • Access Not Enough AI Tutoring 2026 β€” Access is Not Enough: Human Support Improves Engagement with AI Tutoring
  • Adapt Adaptive Lesson Plan Transformer β€” AdaPT: Adaptive Lesson Plan Transformer for Cross-Regional and Differentiated Instruction
  • Agent Voice Accents K12 Group Learning β€” Exploring How Agent Voice Accents Shape Human-AI Collaboration in K-12 Group Learning
  • Agentic AI Education Scoping Review β€” Agentic AI in Education: A Scoping Review of Research Landscape, Capabilities, and the Frontier Agent Paradigm
  • Agentic AI Pedagogical Best Practice 2026 β€” Agentic AI and Pedagogical Best Practice: The Tension Between Automation and Learning
  • Agentic Education Coding β€” Agentic Education with AI Coding Assistants
  • Agentic Literacy Debt β€” Agentic Literacy Debt: A Structural Problem the AI Literacy Field Has Not Yet Named
  • Agents That Teach Incidental Learning β€” Agents That Teach: Designing Incidental Learning Back into AI-Assisted Software Development
  • Agreement Not Quality LLM Coding Verification β€” Agreement Is Not Quality: Blind Expert Verification of Human and LLM Qualitative Coding When Human Consensus Is Not G...
  • AI Agents Constructive Conflict Design Education 2026 β€” Enacting Constructive Conflicts with AI Agents to Enhance Reconsideration among Novice Interaction Designers
  • AI Agents Peer Learning Discourse β€” When AI Agents Teach Each Other: Discourse Patterns Resembling Peer Learning in the Moltbook Community
  • AI Assistance Discretionary Feedback β€” AI Assistance for Discretionary Work: Increasing Feedback Provision in Higher Education
  • AI Assisted Learning Modes Eeg β€” An exploratory behavioral and electroencephalographic study of artificial intelligence-assisted learning modes in hig...
  • AI Availability Student Motivation β€” Why Put in This Much Effort?": How AI Availability Shapes Students’ Motivation in Introductory Programming
  • AI Campus Wellbeing Tools β€” AI-Driven Tools for Enhancing Campus Well-being: Prevention and Intervention
  • AI Changing Teaching Workflows β€” How AI Is Changing Teaching Workflows
  • AI Coaching RL Skill Development β€” AI Coaching for Accelerating Human Skill Development with Reinforcement Learning
  • Citation

    Hazra, R., Ghuku, B., Marchenko, I., Tokarieva, Y., Layek, S., Banerjee, S., Stoyanovich, J., & Pechenizkiy, M. (2026). SafeTutors: Benchmarking Pedagogical Safety in AI Tutoring Systems. arXiv:2603.17373.