Research Article
SafeTutors: Benchmarking Pedagogical Safety in AI Tutoring Systems
Synthesis: Conventional Large Language Models (LLMs) safety benchmarks focus on toxic outputs, jailbreaks, and bias. In education the primary risks are quieter: as Hazra et al. put it, "Solving problems correctly and avoiding toxic language does not make a tutor safe. Tutoring-specific harm is qualitatively different." SafeTutors is a benchmark that jointly evaluates safety and pedagogy in AI tutoring systems across mathematics, physics, and chemistry, and it finds that every tested model shows broad pedagogical harm, with failure rates escalating from 17.7% in single-turn to 77.8% in multi-turn student–tutor dialogue. The harms it measures — answer over-disclosure, misconception reinforcement, abdication of Scaffolding — look "helpful" on the surface: the student gets a correct answer quickly. The long-term effect is learning atrophy, and because these failures worsen as conversations lengthen, single-turn "safe/helpful" results systematically mask tutor failure. Tutoring harm is thus qualitatively different from content harm, and an 11-dimension, 48-sub-risk taxonomy supplies a vocabulary for it grounded in learning science.
Key Findings
- Universal harm across all models. All 11 tested models — 3.8B–72B open-weight systems plus GPT-5-mini — exhibited broad pedagogical harm, and none was consistently safe across subjects or interaction modes.
- Scale is not a fix. Larger models were not reliably safer, and raw helpfulness correlated only weakly with pedagogical safety, challenging the assumption that general capability gains produce better tutors.
- Multi-turn dialogue dramatically worsens behaviour. Harm rates rose from 17.7% in single-turn interactions to 77.8% in multi-turn conversations spanning 5–8 turns, showing that sustained tutoring progressively erodes safety.
- Single-turn evaluation is misleading. Models that appeared safe in one-turn evaluations degraded systematically as conversations extended, so "safe/helpful" single-turn output is not evidence of a safe tutor.
- Harms are subject-dependent. Violation patterns varied significantly across mathematics, physics, and chemistry, indicating that mitigations must be discipline-aware rather than transferred wholesale between science domains.
- An 11-dimension, 48-sub-risk taxonomy grounds the evaluation. SafeTutors derives its risk categories from learning-science literature, tying tutoring failure to cognitive offloading, metacognitive suppression, and diminished learner Learner Agency.
Why Tutoring Safety Is Different
The dominant paradigm evaluates AI tutors on Problem Solving accuracy and generic safety (toxicity, refusal). SafeTutors argues the tutoring-specific risks sit elsewhere:
- Answer over-disclosure — revealing solutions rather than facilitating guided discovery
- Misconception reinforcement — validating or ignoring student misunderstandings
- Abdication of Scaffolding — failing to provide appropriate structured support
- Erosion of productive struggle — short-circuiting the cognitive work that consolidates understanding
Each of these appears benign from the surface: the student gets a correct answer quickly, and nothing toxic was ever said. The damage is deferred, showing up later as learning atrophy and dependence rather than as a violation a content filter could catch. That asymmetry is why a tutor can be technically accurate and "safe" by conventional metrics while systematically undermining learning.
The SafeTutors Risk Taxonomy
Hazra et al. (2026) derive 11 harm dimensions and 48 sub-risks from learning-science literature, giving the evaluation a theoretically grounded vocabulary:
| Dimension | Core Concern | Key Examples |
|---|---|---|
| Cognitive | Interferes with knowledge internalization | Cognitive offloading, fluency illusion, shallow procedural learning |
| Epistemic | Weakens justification/evaluation ability | Unverified authority, source opaqueness, false consensus |
| Metacognitive | Erodes monitoring and self-reflection | External validation dependence, reflection bypass, learned helplessness |
| Motivational-Affective | Undermines curiosity and persistence | Shortcut temptation, performance-over-mastery, emotional disengagement |
| Developmental & Equity | Fails to calibrate to learner level | Cognitive load mismatch, unequal benefit, cultural bias |
| Instructional Alignment | Departs from learning goals | Pedagogical drift, goal misidentification, hidden curriculum |
| Behavioral & Inquiry | Enables shortcuts/dishonesty | Answer-seeking bypass, assignment outsourcing |
| Ethical-Epistemic Integrity | Compromises intellectual ownership | Blurred authorship, misrepresentation of understanding |
| Informational-Semantic | Embeds factual inaccuracies | Fabrication, misleading scientific explanation |
| Reflective-Critical | Suppresses evidence-weighing | Over-smooth acceptance, no metacognitive challenge |
| Pedagogical Relationship | Dysfunctional learner-system dynamic | Over-trust in AI authority, loss of learner agency |
Relationship to Broader Debates
SafeTutors sits at the intersection of AI misuse and learning harm and tutoring-specific evaluation. It sharpens several strands of wiki discussion:
- The LLM Fallacy: Misattribution in AI-Assisted Cognitive Workflows — the fluency illusion (Cognitive dimension) and misrepresentation of understanding (Ethical-Epistemic dimension) are tutoring-specific instantiations of the LLM Fallacy
- Self-Regulated Learning — Motivational-Affective harms undermine the reciprocal loop between self-regulation and Motivation, just as Metacognitive and Reflective-Critical harms suppress monitoring
- Transfer of Learning — Cognitive offloading and shallow procedural learning undermine transfer, and the taxonomy supplies a mechanistic account of why
- Hallucination Risk and Trust — Informational-Semantic failures and over-trust in AI authority compound each other when a fluent tutor is also confidently wrong
Implications for Evaluation, Design, and Policy
- Evaluation: tutor safety must be measured with multi-turn, discipline-specific benchmarks rather than single-turn toxicity screens; behavioral evaluation pipelines should treat multi-turn degradation as the default risk.
- Design: Guardrails must target pedagogical failure modes (over-disclosure, misconception reinforcement) rather than content correctness alone, and the taxonomy can steer training-time approaches such as Pedagogical Safety in Educational Reinforcement Learning and Application-Driven Pedagogical Knowledge Optimization of Open-Source LLMs via Reinforcement Learning and Supervised.
- Policy: procurement criteria for educational AI regulation should require pedagogical safety audits alongside accuracy metrics; systems like EduZone: A Framework for Evaluating LLM Safety for K-12 Students and Teachers and VETTING: A dual-LLM framework for in-loop safety verification via policy isolation in educational AI that prioritize multi-turn safety verification are essential rather than optional.
For K-12 contexts, where the stakes of pedagogical harm are highest and student oversight is thinnest, SafeTutors provides empirical evidence that current models are not safe enough for unsupervised deployment.
What this means for practice
- Designers. Evaluate tutors across multi-turn conversations of 5–8 turns rather than single responses: harm rates rose from 17.7% single-turn to 77.8% multi-turn across all 11 tested models.
- Designers. Do not treat parameter count or proprietary alignment as a safety strategy — the 72B model had a lower harm rate than the 7B on only 17 of 33 single-turn subject–dimension pairs.
- Designers. Target pedagogical failure modes directly — answer over-disclosure, misconception reinforcement, abdication of scaffolding — instead of relying on toxicity and correctness filters.
- Designers. Run subject-specific audits, because violation patterns differed significantly across mathematics, physics, and chemistry, so a mitigation validated in one domain need not transfer.
- Designers. Require multi-turn pedagogical safety evidence before unsupervised K–12 deployment, since harm deepens as conversations lengthen.
Limitations
- The benchmark covers three STEM subjects — mathematics, physics, and chemistry — and 11 models from 3.8B to 72B parameters plus GPT-5-mini, so conclusions stay inside that set.
- Multi-turn evaluation spans 5–8 turns, a shorter horizon than semester-long tutoring, so longer-term degradation is inferred from a capped window.
- Pedagogical scoring is automated using DeepSeek-32B and validated by two doctoral students on a stratified sample of 900 single-turn responses and 300 multi-turn conversations (Cohen's κ = 0.76), leaving most of the output set checked only by the model.
- The 11-dimension, 48-sub-risk taxonomy is derived from learning-science literature and assembled by the authors rather than empirically discovered, so dimension boundaries are analytic choices.
Citation
Hazra, R., Ghuku, B., Marchenko, I., Tokarieva, Y., Layek, S., Banerjee, S., Stoyanovich, J., & Pechenizkiy, M. (2026). SafeTutors: Benchmarking Pedagogical Safety in AI Tutoring Systems.