AI Tutor Safety and Pedagogical Harms

Created: 2026-05-07 | Tags: pedagogical-safetyintelligent-tutoringadaptive-learningk-12higher-edllmbias-mitigation
πŸ“„ Full text: arXiv:2603.17373 Β· local
"Solving problems correctly and avoiding toxic language does not make a tutor safe. Tutoring-specific harm is qualitatively different." SafeTutors exposes that all tested models show broad pedagogical harm, with failures escalating from 17.7% in single-turn to 77.8% in multi-turn student-tutor dialogue.^hazra-safetutors-pedagogical-safety-2026

Why Tutoring Safety Is Different

Conventional LLM safety benchmarks focus on toxic outputs, jailbreaks, and bias. In education, the primary risks are quieter:

These harms appear "helpful" to surface inspection: the student gets a correct answer quickly. But the long-term effect is learning atrophy.

The SafeTutors Risk Taxonomy

Hazra et al. (2026) derive 11 harm dimensions and 48 sub-risks from learning-science literature:

Dimension Core Concern Key Examples
Cognitive Interferes with knowledge internalization Cognitive offloading, fluency illusion, shallow procedural learning
Epistemic Weakens justification/evaluation ability Unverified authority, source opaqueness, false consensus
Metacognitive Erods monitoring and self-reflection External validation dependence, reflection bypass, learned helplessness
Motivational-Affective Undermines curiosity and persistence Shortcut temptation, performance-over-mastery, emotional disengagement
Developmental & Equity Fails to calibrate to learner level Cognitive load mismatch, unequal benefit, cultural bias
Instructional Alignment Departs from learning goals Pedagogical drift, goal misidentification, hidden curriculum
Behavioral & Inquiry Enables shortcuts/dishonesty Answer-seeking bypass, assignment outsourcing
Ethical-Epistemic Integrity Compromises intellectual ownership Blurred authorship, misrepresentation of understanding
Informational-Semantic Embeds factual inaccuracies Fabrication, misleading scientific explanation
Reflective-Critical Suppresses evidence-weighing Over-smooth acceptance, no metacognitive challenge
Pedagogical Relationship Dysfunctional learner-system dynamic Over-trust in AI authority, loss of learner agency

Critical Findings

1. Universal harm: All 11 tested models (3.8B–72B open-weight + GPT-5-mini) exhibited broad pedagogical harm 2. Scale is not a fix: Larger models were not reliably safer; raw helpfulness correlates weakly with pedagogical safety 3. Multi-turn degradation: Harm rates rose from 17.7% (single-turn) to 77.8% (multi-turn), showing that sustained tutoring interaction progressively erodes safety 4. Discipline-aware mitigations needed: Harms varied significantly across math, physics, and chemistry 5. Single-turn evaluation is misleading: "Safe" single-turn responses masked systematic failure when conversations extended to 5–8 turns

Relationship to Broader Debates

Implications

Related Pages


πŸ“Ž 1 other page tagged ai-tutor-safety-harms