On this page

Synthesis: Conventional Large Language Models (LLMs) safety benchmarks focus on toxic outputs, jailbreaks, and bias. In education the primary risks are quieter: as Hazra et al. put it, "Solving problems correctly and avoiding toxic language does not make a tutor safe. Tutoring-specific harm is qualitatively different." SafeTutors is a benchmark that jointly evaluates safety and pedagogy in AI tutoring systems across mathematics, physics, and chemistry, and it finds that every tested model shows broad pedagogical harm, with failure rates escalating from 17.7% in single-turn to 77.8% in multi-turn student–tutor dialogue. The harms it measures — answer over-disclosure, misconception reinforcement, abdication of Scaffolding — look "helpful" on the surface: the student gets a correct answer quickly. The long-term effect is learning atrophy, and because these failures worsen as conversations lengthen, single-turn "safe/helpful" results systematically mask tutor failure. Tutoring harm is thus qualitatively different from content harm, and an 11-dimension, 48-sub-risk taxonomy supplies a vocabulary for it grounded in learning science.

Key Findings

  1. Universal harm across all models. All 11 tested models — 3.8B–72B open-weight systems plus GPT-5-mini — exhibited broad pedagogical harm, and none was consistently safe across subjects or interaction modes.
  2. Scale is not a fix. Larger models were not reliably safer, and raw helpfulness correlated only weakly with pedagogical safety, challenging the assumption that general capability gains produce better tutors.
  3. Multi-turn dialogue dramatically worsens behaviour. Harm rates rose from 17.7% in single-turn interactions to 77.8% in multi-turn conversations spanning 5–8 turns, showing that sustained tutoring progressively erodes safety.
  4. Single-turn evaluation is misleading. Models that appeared safe in one-turn evaluations degraded systematically as conversations extended, so "safe/helpful" single-turn output is not evidence of a safe tutor.
  5. Harms are subject-dependent. Violation patterns varied significantly across mathematics, physics, and chemistry, indicating that mitigations must be discipline-aware rather than transferred wholesale between science domains.
  6. An 11-dimension, 48-sub-risk taxonomy grounds the evaluation. SafeTutors derives its risk categories from learning-science literature, tying tutoring failure to cognitive offloading, metacognitive suppression, and diminished learner Learner Agency.

Why Tutoring Safety Is Different

The dominant paradigm evaluates AI tutors on Problem Solving accuracy and generic safety (toxicity, refusal). SafeTutors argues the tutoring-specific risks sit elsewhere:

  • Answer over-disclosure — revealing solutions rather than facilitating guided discovery
  • Misconception reinforcement — validating or ignoring student misunderstandings
  • Abdication of Scaffolding — failing to provide appropriate structured support
  • Erosion of productive struggle — short-circuiting the cognitive work that consolidates understanding

Each of these appears benign from the surface: the student gets a correct answer quickly, and nothing toxic was ever said. The damage is deferred, showing up later as learning atrophy and dependence rather than as a violation a content filter could catch. That asymmetry is why a tutor can be technically accurate and "safe" by conventional metrics while systematically undermining learning.

The SafeTutors Risk Taxonomy

Hazra et al. (2026) derive 11 harm dimensions and 48 sub-risks from learning-science literature, giving the evaluation a theoretically grounded vocabulary:

Dimension Core Concern Key Examples
Cognitive Interferes with knowledge internalization Cognitive offloading, fluency illusion, shallow procedural learning
Epistemic Weakens justification/evaluation ability Unverified authority, source opaqueness, false consensus
Metacognitive Erodes monitoring and self-reflection External validation dependence, reflection bypass, learned helplessness
Motivational-Affective Undermines curiosity and persistence Shortcut temptation, performance-over-mastery, emotional disengagement
Developmental & Equity Fails to calibrate to learner level Cognitive load mismatch, unequal benefit, cultural bias
Instructional Alignment Departs from learning goals Pedagogical drift, goal misidentification, hidden curriculum
Behavioral & Inquiry Enables shortcuts/dishonesty Answer-seeking bypass, assignment outsourcing
Ethical-Epistemic Integrity Compromises intellectual ownership Blurred authorship, misrepresentation of understanding
Informational-Semantic Embeds factual inaccuracies Fabrication, misleading scientific explanation
Reflective-Critical Suppresses evidence-weighing Over-smooth acceptance, no metacognitive challenge
Pedagogical Relationship Dysfunctional learner-system dynamic Over-trust in AI authority, loss of learner agency

Relationship to Broader Debates

SafeTutors sits at the intersection of AI misuse and learning harm and tutoring-specific evaluation. It sharpens several strands of wiki discussion:

Implications for Evaluation, Design, and Policy

For K-12 contexts, where the stakes of pedagogical harm are highest and student oversight is thinnest, SafeTutors provides empirical evidence that current models are not safe enough for unsupervised deployment.

What this means for practice

  • Designers. Evaluate tutors across multi-turn conversations of 5–8 turns rather than single responses: harm rates rose from 17.7% single-turn to 77.8% multi-turn across all 11 tested models.
  • Designers. Do not treat parameter count or proprietary alignment as a safety strategy — the 72B model had a lower harm rate than the 7B on only 17 of 33 single-turn subject–dimension pairs.
  • Designers. Target pedagogical failure modes directly — answer over-disclosure, misconception reinforcement, abdication of scaffolding — instead of relying on toxicity and correctness filters.
  • Designers. Run subject-specific audits, because violation patterns differed significantly across mathematics, physics, and chemistry, so a mitigation validated in one domain need not transfer.
  • Designers. Require multi-turn pedagogical safety evidence before unsupervised K–12 deployment, since harm deepens as conversations lengthen.

Limitations

  • The benchmark covers three STEM subjects — mathematics, physics, and chemistry — and 11 models from 3.8B to 72B parameters plus GPT-5-mini, so conclusions stay inside that set.
  • Multi-turn evaluation spans 5–8 turns, a shorter horizon than semester-long tutoring, so longer-term degradation is inferred from a capped window.
  • Pedagogical scoring is automated using DeepSeek-32B and validated by two doctoral students on a stratified sample of 900 single-turn responses and 300 multi-turn conversations (Cohen's κ = 0.76), leaving most of the output set checked only by the model.
  • The 11-dimension, 48-sub-risk taxonomy is derived from learning-science literature and assembled by the authors rather than empirically discovered, so dimension boundaries are analytic choices.

Citation

Hazra, R., Ghuku, B., Marchenko, I., Tokarieva, Y., Layek, S., Banerjee, S., Stoyanovich, J., & Pechenizkiy, M. (2026). SafeTutors: Benchmarking Pedagogical Safety in AI Tutoring Systems.

Embed this page

Copy the code below to embed a chromeless version of this page in a learning management system or other website. The embedded view hides the site header, navigation, and footer.