Research Article
Simulating Students or Sycophantic Problem Solving? On Misconception Faithfulness of LLM Simulators
Synthesis: This paper exposes a critical failure mode in using LLMs as simulated students for Intelligent Tutoring development and evaluation. The authors introduce misconception faithfulness — the property that a simulated student holds a coherent, misconception-driven belief state and updates it only when feedback addresses the underlying misconception — and show that across seven LLMs (4B to 120B parameters), simulators exhibit near-zero faithfulness. The core finding is a sycophantic failure mode: when given any corrective signal, Large Language Models (LLMs) simulators abandon their assigned misconception persona and re-solve the problem from internal knowledge. They behave as problem-solvers, not as students with stable Misconceptions about AI. Using the novel Selective Flip Score (SFS), the authors quantify this: simulators flip their answers at similarly high rates regardless of whether feedback is targeted, misaligned, or generic. This connects directly to Sycophancy is an Educational Safety Risk: Why LLM Tutors Need Sycophancy Benchmarks, which identified sycophancy as an educational safety risk in LLM tutors. Here the sycophancy is inverted: simulated students capitulate to feedback rather than maintaining authentic misconception-driven behavior. Both papers together establish sycophancy as a bidirectional problem in AIED — affecting both tutor and student roles. The post-training pipeline — combining supervised fine-tuning, preference optimization, and RL with SFS-aligned rewards — achieved SFS gains up to +0.56, demonstrating that misconception faithfulness is trainable. This has implications for SafeTutors: Benchmarking Pedagogical Safety in AI Tutoring Systems: if student simulators used for tutor safety testing are themselves unfaithful, safety evaluations conducted on them may systematically miss harm patterns that real students would exhibit. For Student Experience and Benchmark development, this paper motivates a paradigm shift from static output matching toward interactive, belief-aware student modeling — a theme that also resonates with PersonaVLM: Long-Term Personalized Multimodal LLMs and the behavioral evaluation framework in The Missing Evaluation Axis: What 10,000 Student Submissions Reveal About AI Tutor Effectiveness.
What this means for practice
- Software developers. Never validate a student simulator on output similarity alone. Across seven instruction-tuned models from 4B to 120B parameters, simulators flipped to the correct answer at near-uniform rates under targeted, misaligned, and generic feedback, so plausible student-like answers say nothing about the belief state behind them.
- Software developers. Score simulators with an interactive, misconception-contrastive metric such as the Selective Flip Score, and train against it. Preference optimization alone underperformed; supervised fine-tuning plus preference optimization plus RL with SFS-aligned rewards produced gains up to +0.56.
- Software developers. Test simulator faithfulness before trusting a tutor evaluation built on it. If simulators capitulate to any corrective signal, safety and quality testing on them will systematically miss the harm patterns real students would trigger.
- Learners. When an AI tutor tells you an answer is wrong, ask what was wrong with your reasoning. The simulator behavior documented here shows how easily a correction changes the answer without touching the underlying misunderstanding.
Limitations
- Both the teacher feedback and the response classifications were generated by LLMs (GPT-OSS-120B) rather than human raters, a choice the authors made for scalable, reproducible evaluation across many models.
- The framework covers mathematics domains with misconceptions drawn from structured inventories — Malrule arithmetic error patterns and Eedi multiple-choice distractors — and the authors state that less structured domains remain future work.
- Dataset composition is uneven: judge-filtered synthetic items made up 0.2% of Malrule but 14.6% of Eedi, so dataset-level differences can confound cross-model comparisons.
- Model coverage is narrow relative to the claim: seven instruction-tuned LLMs from the Llama, Qwen3, and GPT-OSS families (4B–120B), evaluated on small splits such as Malrule's 790 training and 210 test items and Eedi's 800 test items.
Citation
Do, H., Sonkar, S., & Sachan, M. (2026). Simulating Students or Sycophantic Problem Solving? On Misconception Faithfulness of LLM Simulators.