🧠 AI Ed Wiki

Conventional LLM safety benchmarks focus on toxic outputs, jailbreaks, and bias. In education, the primary risks are quieter:

"Solving problems correctly and avoiding toxic language does not make a tutor safe. Tutoring-specific harm is qualitatively different." SafeTutors exposes that all tested models show broad pedagogical harm, with failures escalating from 17.7% in single-turn to 77.8% in multi-turn student-tutor dialogue.^Hazra Safetutors Pedagogical Safety 2026

Why Tutoring Safety Is Different

Conventional LLM safety benchmarks focus on toxic outputs, jailbreaks, and bias. In education, the primary risks are quieter:

  • Answer over-disclosure — Revealing solutions rather than facilitating guided discovery
  • Misconception reinforcement — Validating or ignoring student misunderstandings
  • Abdication of scaffolding — Failing to provide appropriate structured support
  • Erosion of productive struggle — Short-circuiting the cognitive work that consolidates understanding
  • These harms appear "helpful" to surface inspection: the student gets a correct answer quickly. But the long-term effect is learning atrophy.

    The SafeTutors Risk Taxonomy

    Hazra et al. (2026) derive 11 harm dimensions and 48 sub-risks from learning-science literature:

    DimensionCore ConcernKey Examples
    CognitiveInterferes with knowledge internalizationCognitive offloading, fluency illusion, shallow procedural learning
    EpistemicWeakens justification/evaluation abilityUnverified authority, source opaqueness, false consensus
    MetacognitiveErods monitoring and self-reflectionExternal validation dependence, reflection bypass, learned helplessness
    Motivational-AffectiveUndermines curiosity and persistenceShortcut temptation, performance-over-mastery, emotional disengagement
    Developmental & EquityFails to calibrate to learner levelCognitive load mismatch, unequal benefit, cultural bias
    Instructional AlignmentDeparts from learning goalsPedagogical drift, goal misidentification, hidden curriculum
    Behavioral & InquiryEnables shortcuts/dishonestyAnswer-seeking bypass, assignment outsourcing
    Ethical-Epistemic IntegrityCompromises intellectual ownershipBlurred authorship, misrepresentation of understanding
    Informational-SemanticEmbeds factual inaccuraciesFabrication, misleading scientific explanation
    Reflective-CriticalSuppresses evidence-weighingOver-smooth acceptance, no metacognitive challenge
    Pedagogical RelationshipDysfunctional learner-system dynamicOver-trust in AI authority, loss of learner agency

    Critical Findings

    1. Universal harm: All 11 tested models (3.8B–72B open-weight + GPT-5-mini) exhibited broad pedagogical harm

    2. Scale is not a fix: Larger models were not reliably safer; raw helpfulness correlates weakly with pedagogical safety

    3. Multi-turn degradation: Harm rates rose from 17.7% (single-turn) to 77.8% (multi-turn), showing that sustained tutoring interaction progressively erodes safety

    4. Discipline-aware mitigations needed: Harms varied significantly across math, physics, and chemistry

    5. Single-turn evaluation is misleading: "Safe" single-turn responses masked systematic failure when conversations extended to 5–8 turns

    Relationship to Broader Debates

  • Tutoring Specific Vs General AI — SafeTutors reveals that even "helpful" general-purpose AI produces systematic tutoring harm; pedagogical design is not an add-on but a safety requirement
  • Metacognition — The Metacognitive and Reflective-Critical dimensions directly map to metacognitive suppression risks
  • Self Regulated Learning — Motivational-Affective harms undermine the SRL↔motivation reciprocal loop
  • Transfer Of Learning — Cognitive offloading and shallow learning directly undermine transfer; SafeTutors provides a mechanistic taxonomy for why
  • LLM Fallacy Misattribution — Fluency illusion (Cognitive dimension) and misrepresentation of understanding (Ethical-Epistemic dimension) are tutoring-specific instantiations of the LLM Fallacy
  • Implications

  • Evaluation: Tutor safety must be measured with multi-turn, discipline-specific benchmarks, not single-turn toxicity screens
  • Design: Guardrails must target pedagogical failure modes (over-disclosure, misconception reinforcement) not just content correctness
  • Policy: Procurement criteria for educational AI should include pedagogical safety audits alongside accuracy metrics
  • Connected Concepts

  • Metacognition
  • Self Regulated Learning
  • Connected Articles

  • Hazra Safetutors Pedagogical Safety 2026
  • Tutoring Specific Vs General AI
  • Transfer Of Learning
  • LLM Fallacy Misattribution
  • Citation

    Hazra, R., Ghuku, B., Marchenko, I., Tokarieva, Y., Layek, S., Banerjee, S., Stoyanovich, J., & Pechenizkiy, M. (2026). SafeTutors: Benchmarking Pedagogical Safety in AI Tutoring Systems. arXiv:2603.17373.