Research Article
LLM-Based Educational Simulation: Evaluating Temporal Student Persona Stability Across ADHD Profiles
Synthesis: Gonnermann-Müller, Haase & Leins (2026) evaluate whether LLM-generated student personas simulating ADHD profiles maintain stable and realistic behavioral patterns over time. This addresses a critical question for using LLMs in educational research and teacher training: can simulated learners reliably represent neurodivergent students?
Why This Matters
Using LLMs to simulate students is an emerging practice in educational research, but the temporal stability of these simulations — especially for neurodivergent profiles — has been underexamined. If LLM-generated personas drift or become inconsistent, they cannot serve as valid proxies for real students in:
- Teacher training simulations
- Adaptive Learning testing
- Intelligent Tutoring system evaluation
- Learning Analytics research methodology
Connections to Knowledge Base
This work extends the PersonaVLM: Long-Term Personalized Multimodal LLMs discourse on how LLMs represent learners over time, but applies it to simulation validity rather than tutoring personalization. The focus on ADHD profiles connects to broader Student Experience research and highlights gaps in The Evidence Base on AI in K-12: A 2026 Review — the Stanford SCALE review found few studies with adequate causal inference for special education populations.
The simulation methodology also raises questions about SafeTutors: Benchmarking Pedagogical Safety in AI Tutoring Systems — if tutoring systems are tested on simulated neurodivergent learners, do the safety assessments generalize? This echoes The Aftermath of DrawEduMath: Vision Language Models Underperform with Struggling Students and Misdiagnose Errors concerns about AI systems that underperform with specific student populations.
Open Questions
- How do LLM-simulated ADHD profiles compare to Multimodal Dialogue in STEM Education systems that work with real neurodivergent students?
- Can temporal stability be improved through prompt engineering or fine-tuning?
- What is the ethical boundary for using simulated students in RCT designs?
What this means for practice
- Software developers. Build scripted, task-anchored interactions rather than open-ended chat when simulated learners must stay in character: scripted interactions with explicit task prompts eliminated observer-rated behavioral drift entirely, a reduction of up to 97% relative to unscripted dialog.
- Software developers. Treat interaction structure as a stronger lever than model selection. The study crossed five LLMs with three prompt designs and four persona conditions, and stability turned out to be conditional on interaction design rather than intrinsic LLM capability.
- Software developers. Test the middle of your persona distribution, not only the extremes: within-conversation drift occurred in unscripted dialog for both high- and moderate-intensity ADHD personas, so partially specified profiles are the ones simulated worst.
- Learners. Use simulated student scenarios as rehearsal, not as a model of real neurodivergent classmates: self-reported persona characteristics stayed stable while observer-rated behavioral expression of high- and moderate-intensity personas declined across the 9-turn conversations.
- Software developers. Set the persona explicitly, because without persona instructions baseline LLM student representation skews toward high ADHD symptoms — a bias the authors trace to pretraining material drawn disproportionately from clinical and special education contexts.
Limitations
- The study covers one diagnostic construct only; generalization to personas combining multiple human characteristics, and to comorbid conditions such as ADHD with anxiety, remains for future research, and the authors note this limits ecological validity for educational simulation.
- Only two interaction structures were tested — scripted and unscripted. Other structures such as increasing Scaffolding, or supportive versus challenging tutor personas, were not explored, so the minimal intervention for behavioral stability is unknown.
- Observer-rated behavioral expression was scored by three independent LLM raters blind to persona instructions, so the behavioral stability measures are themselves model-generated rather than human-coded.
- Between-conversation stability rests on single-turn, context-free instantiations (N = 4,968), and within-conversation stability on 20 conversations of 9 turns (N = 3,952) — a short interaction window relative to the sustained, path-dependent interactions of real tutoring or teacher training.
Citation
Gonnermann-Müller, J., Haase, J., & Leins, N. (2026). LLM-Based Educational Simulation: Evaluating Temporal Student Persona Stability Across ADHD Profiles.