🧠 AI Ed Wiki

Synthesis: Semi-structured interviews with 12 teachers who tutored LLM-simulated students (MathDial dataset) reveal key authenticity gaps: overly complex language, lack of emotions, unnatural attentiveness, and logical inconsistency. The study categorizes four real-world student behavior types along scaffolding and presence dimensions, and provides design guidelines for building higher-fidelity LLM student simulations.

Methodology

Martynova et al. interviewed 12 teachers who had extensively interacted with LLM-simulated students during collection of the MathDial dialogue tutoring dataset. The study used a mixed-method approach grounded in two frameworks:

  • Community of Inquiry (CoI) β€” capturing social and cognitive presence in learning interactions
  • Scaffolding theory β€” effective teaching through graduated support
  • Teachers tutored LLM students in K-12 math problem-solving dialogues, then rated realism and described deviations from authentic student behavior.

    Key Findings

    Authenticity Gaps in LLM Students

    IssueDescription
    Language complexityResponses too technical, lengthy, and formal for K-12 students
    Emotional absenceLack of frustration, fear, embarrassment, or disengagement
    Unnatural attentivenessStudents too engaged; never lose focus or go silent
    Logical inconsistencyKnowledge jumps without gradual building; no forgetting
    No question-askingTeachers had too much control over discussion flow

    Four Student Behavior Categories

    The study classifies real-world student behaviors along two dimensions:

    High Scaffolding NeedsLow Scaffolding Needs
    Social PresenceShort/simple writing, negative emotions, disengagementAsking questions, disagreeing with teacher
    Cognitive PresenceGradual knowledge-building, memory/forgettingChanging tactics based on feedback

    LLMs captured the bottom-right quadrant reasonably well but failed to represent the other three categories.

    Design Guidelines

    1. Diverse personalities β€” model Big Five personality traits to produce varied engagement levels and emotional responses

    2. Gradual knowledge building β€” integrate knowledge tracing to avoid unrealistic knowledge jumps

    3. Model forgetting β€” account for memory decay over time

    4. Promote question-asking β€” use context-aware triggers for the LLM student to ask questions

    5. Vary language complexity β€” regulate response length, formality, and introduce age-appropriate errors

    6. Allow disengagement β€” let simulated students lose focus or stay silent, providing authentic teaching challenges

    Significance

  • Teacher training: more realistic LLM student simulations enable scalable practice for pre-service and in-service teachers
  • Validation gap: only 3% of studies simulating learners do post-factum validation β€” this study provides a framework for it
  • MathDial is the only publicly available dataset of real teacher/LLM-student interactions
  • Addresses the growing trend of using unvalidated LLM simulations in educational contexts
  • Connected Concepts

  • K 12
  • Knowledge Tracing
  • LLM
  • Scaffolding
  • Connected Articles

  • Aaai2026 Prompting Literacy K12 β€” Learning to Use AI for Learning: Teaching Responsible Use of AI Chatbot to K-12 Students Through an AI Literacy Module
  • Academiclaw Student Agent Benchmark β€” AcademiClaw: When Students Set Challenges for AI Agents
  • Access Not Enough AI Tutoring 2026 β€” Access is Not Enough: Human Support Improves Engagement with AI Tutoring
  • Adapt Adaptive Lesson Plan Transformer β€” AdaPT: Adaptive Lesson Plan Transformer for Cross-Regional and Differentiated Instruction
  • Agent Voice Accents K12 Group Learning β€” Exploring How Agent Voice Accents Shape Human-AI Collaboration in K-12 Group Learning
  • Agentic AI Education Scoping Review β€” Agentic AI in Education: A Scoping Review of Research Landscape, Capabilities, and the Frontier Agent Paradigm
  • Agentic AI Pedagogical Best Practice 2026 β€” Agentic AI and Pedagogical Best Practice: The Tension Between Automation and Learning
  • Agentic Education Coding β€” Agentic Education with AI Coding Assistants
  • Agentic Literacy Debt β€” Agentic Literacy Debt: A Structural Problem the AI Literacy Field Has Not Yet Named
  • Agents That Teach Incidental Learning β€” Agents That Teach: Designing Incidental Learning Back into AI-Assisted Software Development
  • Agreement Not Quality LLM Coding Verification β€” Agreement Is Not Quality: Blind Expert Verification of Human and LLM Qualitative Coding When Human Consensus Is Not G...
  • AI Agents Constructive Conflict Design Education 2026 β€” Enacting Constructive Conflicts with AI Agents to Enhance Reconsideration among Novice Interaction Designers
  • AI Agents Peer Learning Discourse β€” When AI Agents Teach Each Other: Discourse Patterns Resembling Peer Learning in the Moltbook Community
  • AI Assistance Discretionary Feedback β€” AI Assistance for Discretionary Work: Increasing Feedback Provision in Higher Education
  • AI Assisted Learning Modes Eeg β€” An exploratory behavioral and electroencephalographic study of artificial intelligence-assisted learning modes in hig...
  • AI Availability Student Motivation β€” Why Put in This Much Effort?": How AI Availability Shapes Students’ Motivation in Introductory Programming
  • AI Campus Wellbeing Tools β€” AI-Driven Tools for Enhancing Campus Well-being: Prevention and Intervention
  • AI Changing Teaching Workflows β€” How AI Is Changing Teaching Workflows
  • AI Coaching RL Skill Development β€” AI Coaching for Accelerating Human Skill Development with Reinforcement Learning
  • AI Education Global Capacity β€” What AI in Education Needs Next: Lessons from Youth Leaders Across Five Countries
  • AI Enabled Serious Games β€” AI-Enabled Serious Games: Integrating Intelligence and Adaptivity in Training Systems
  • AI Engineering Education Balancing Act β€” Using AI in engineering education: a balancing act, driven by clear purpose
  • AI Generated Traces Novice Programmers β€” AI-Generated Traces for Novice Programmers: Learning Effects and Learner Differences in a Multi-Institutional Study
  • AI In The Wild College β€” AI in the Wild: A Large Scale Analysis of Authentic Interactions of College Students with Generative AI
  • AI Interlocutor L2 Spoken Dialogue β€” What Changes When the Interlocutor Is an AI? Interactional Fluency and Linguistic Uptake in L2 Spoken Dialogue
  • Citation

    Learners?, C.L.E.S.H., Students, T.I.F.T.L., Daheim1,2, D.M.J.M.N., Sachan1, Γ–.N.Y.X.Z.M., Fraser, E.Z.T.D.S., many, L.L.M.O., & aims, F.B.H.L.S.A.P.U.T.S. (2026). Can LLMs Effectively Simulate Human Learners? Teachers' Insights from Tutoring LLM Students. Innovative Use of NLP for Building Educational Applications) DOI: https://aclanthology