On this page

Synthesis: Marquez-Carpintero, Lopez-Sellers & Cazorla (2025) present a thematic review of empirical and methodological studies using LLMs to simulate student behavior in education. They synthesize evidence on how LLM-based agents emulate learner archetypes, respond to instructional inputs, and interact in multi-agent classroom scenarios, and examine implications for curriculum development, instructional evaluation, and teacher training — while flagging persistent concerns around algorithmic bias, evaluation reliability, and alignment with educational objectives. The review frames simulated students as a valuable methodological tool for evaluating pedagogy and modeling diverse learner profiles, and identifies it as the first literature review dedicated specifically to simulating student roles and mechanisms with LLMs.

Key Findings

  1. LLM-based simulated students extend the pre-LLM lineage of symbolic and probabilistic student models — from SCHOLAR and Anderson's LISP Tutor (ACT-R based) through Bayesian Knowledge Tracing — by adding high linguistic realism and behavioral adaptability.
  2. Architectural advances rest on multi-agent systems and memory mechanisms: short- and long-term stores plus iterative reflection (e.g. Classroom Simulacra's Transferable Iterative Reflection, Agent4Edu, EduAgent) that let agents learn from incorrect predictions and maintain coherent learning trajectories.
  3. Student knowledge is modeled through three complementary strategies — direct prompt-based simulation, deep-learning knowledge tracing (DKT, DKVMN, AKT, LLM-KT), and knowledge graphs with heuristics — each trading interpretability against scalability.
  4. Personality frameworks such as the Big Five (and the more contested MBTI) are embedded to simulate learner diversity, support teacher training, and validate differentiated instructional strategies.
  5. Persistent concerns remain around algorithmic bias, evaluation reliability, and alignment with educational objectives, alongside technical and methodological gaps that define the open research agenda.

Background: from rule-based tutors to generative agents

The Simulated Student paradigm predates LLMs. In the 1990–2019 era it centered on internal student representations for adaptive instruction within Intelligent Tutoring Systems, beginning with symbolic, rule-based systems such as SCHOLAR, which used semantic nets to track known concepts. The most ambitious goal was cognitive modeling, exemplified by Anderson's LISP Tutor on the ACT-R architecture, which used model-tracing to diagnose a learner's trajectory against an ideal student model. While offering high fidelity and explainability, these systems suffered from costly knowledge engineering and poor scalability.

The field then shifted toward probabilistic and data-driven methods. Bayesian Knowledge Tracing (BKT) became the canonical approach, using a hidden Markov model to estimate skill mastery over time, followed by analytics-based approaches — Educational Data Mining and Learning Analytics — that applied supervised machine learning to predict outcomes and flag at-risk students. A systematic review of the 2010–2019 literature (Käser & Alexandron, 2024) found that these models represented only "narrow aspects of student learning" and that almost half of simulated-learner studies lacked formal validation — a crisis of fidelity and validation that frames the appeal of LLM-based simulation.

Educational psychology foundations

The review grounds LLM-based simulation in learning theory. From the Constructivism perspective, learning is an active process of knowledge construction, and LLM models incorporate these principles by categorizing their own learning capabilities relative to varying prior knowledge. The Zone of Proximal Development (ZPD) is used to design adaptive interventions within in-context learning and fine-tuning, where demonstrations serve as Scaffolding graded to the model's "cognitive state." Cognitive theories — such as information processing theory — offer tools for modeling metacognitive processes, attention, and cognitive load, which have begun to be integrated into architectures like Agent4Edu and Classroom Simulacra.

Architecture and implementation mechanisms

Recent work optimizes architectural design through multi-agent systems, modular components, and the integration of diverse generative agents. Classroom Simulacra introduces a Transferable Iterative Reflection (TIR) module: during training a reflective agent issues an initial prediction, compares it with ground truth, writes a reflection, and hands it to a novice agent that re-predicts; the loop repeats until accuracy plateaus and the most useful reflections are stored. During testing the model retrieves stored reflections, combines them with the student's history, and predicts future answers. Course knowledge arrives via lecture slides as external stimuli, while prior knowledge acts as internal stimuli. Agent4Edu combines predefined student profiles, memory modules, and action modules for personalized learning scenarios, while SimClass coordinates Teacher, Assistant, and Classmate Agents with defined personality traits to simulate peer dynamics.

Memory management

Memory is a core component of realistic simulation. Agent4Edu's memory system performs three operations: retrieval (extracting pertinent information from short- and long-term stores), writing (promoting raw observations into longer-term memory upon reinforcement), and reflection (summary and corrective forms occurring in long-term memory). Classroom Simulacra's reflection database acts as a form of long-term memory that directly influences future decisions. EduAgent instantiates agents with personality profiles that simulate learning slide by slide, with ablation studies showing that removing past cognitive states from memory significantly reduces performance — highlighting the critical role of contextual memory. MathVC implements short-term memory via dynamic dialogue history and long-term memory via a symbolic character schema. This mirrors general generative-agent work (Park et al., 2023) where behavior is conditioned on a continuously updated memory history.

Knowledge modeling strategies

The review identifies three differentiated strategies for modeling student knowledge:

  • Direct prompt-based simulation. Knowledge levels, mistakes, or learning styles are defined directly in the LLM prompt. This is the most widely employed approach (e.g. EduAgent, TeachTune's Interpret–Reflect–Respond pipeline). Studies applying Item Response Theory show that while some LLMs perform well, they exhibit narrower zones of ZPD; combining human and synthetic responses at a 1:1 ratio yields more accurate item calibration. Integrating retrieval-augmented generation with topic-specific documents is proposed to refine proficiency definition.
  • Knowledge tracing. KT tracks a student's knowledge state over time using knowledge components (KCs). Deep-learning models (DKT, DKVMN, graph-based GNN models, AKT) and LLM-hybrid models such as LLM-KT underpin this strategy. By shifting the task from predicting a correct/incorrect label to generating a coherent response conditioned on an internal knowledge state, KT models evolve into generative student simulators.
  • Knowledge graphs and heuristics. Structured representations organize pedagogical concepts and relationships that agents can query, with heuristics treated as KCs injected into prompts (mastered/confused/unknown). Systems like KnowLearn construct knowledge graphs from real pedagogical data using a heterogeneous graph attention network (HAN) to identify which factors most influence progress.

Personality and individual traits

Educational models traditionally focused on cognitive variables such as knowledge level, overlooking non-cognitive elements. The review examines how psychological frameworks such as the Big Five (Five Factor Model) and MBTI are embedded into LLM agents to simulate distinct behavioral patterns. TeachTune applies the Big Five to Pedagogical Conversational Agents, finding that teachers identify personality, Motivation, and self-regulated learning strategies as key to personalized instruction. The Big Five for Tutoring Conversation (BF-TC) reformulates trait items for educational dialogue and shows strong correspondence with the standard Big Five Inventory. The MBTI, while structured and occasionally integrated, is flagged as controversial, lacking strong empirical support. These personality-enabled agents (e.g. SENEM-AI, TutorUp) help prospective educators recognize and respond to diverse classroom situations and practice personalized feedback without ethical risk.

Evaluation and applications

Simulated students offer a low-risk means of experimenting with pedagogical strategies, curriculum design, and assessment methods before implementation in real classrooms, enhancing scalability and pedagogical safety. Applications span curriculum development, instructional evaluation, and teacher training. Concerns persist, however: the degree to which LLMs faithfully replicate human cognitive and affective processes is under investigation, with ongoing issues of algorithmic bias, limitations in open-access training datasets, and the risk of generating overly idealised or homogenized behaviors. The reliability of a simulated student's fit to a real learner is itself hard to validate.

What this means for practice

  • Instructors. Use simulated students for the rehearsals that are hard or unsafe to run with real learners — calibrating assessment items, trying out tutoring moves, and letting novice teachers respond to difficult classroom situations without exposing pupils.
  • Instructors. Treat simulation output as an analytical adjunct rather than a verdict: the review places current fidelity in an intermediate band, adequate for proof-of-concept and low-stakes analysis but not a substitute for observing authentic students in high-stakes settings.
  • Instructors. Write the persona, not just the prompt: encode knowledge components, trait profiles and illustrative rules, since prompt design is what decides how faithfully an agent behaves, and retrieval augmentation can sharpen proficiency modeling further.
  • Faculty developers. Mix human and synthetic responses instead of choosing between them — in the Item Response Theory studies the review covers, combining real and synthetic responses at a 1:1 ratio produced more accurate item calibration than synthetic data alone.
  • Researchers. Make formal validation the first design decision rather than the last: the pre-LLM systematic review the authors cite found almost half of simulated-learner studies lacked it, and the review reports the same gap persisting, along with evaluation criteria that remain costly and subjective.

Limitations

  • The review is a qualitative thematic synthesis, not a quantitative one: no effect sizes are pooled, and inclusion prioritized peer-reviewed work and significant preprints published between 2021 and 2025, with historically relevant articles added manually outside the stated criteria.
  • Screening began with 808 titles after 408 duplicates were removed, judged by two independent reviewers, and the authors list three inherent limits of the search: possible overrepresentation of studies with positive results, restriction to English-language publications, and the manual handling of Google Scholar.
  • Gray literature was excluded outright — master's and doctoral theses, unpublished institutional documents, developer technical blogs and community-of-practice content — so the corpus reflects peer-reviewed and preprint literature only.
  • The authors state that the field itself lacks agreed validation criteria (dataset comparison, expert judgment and Turing-style tests being costly, subjective and hard to scale) and that behavioral fidelity remains constrained by idealized answers and weak long-term consistency, with most reviewed simulations confined to laboratory settings rather than deployed platforms.

Citation

Marquez-Carpintero, L., Lopez-Sellers, A., & Cazorla, M. (2025). Simulating Students with Large Language Models: A Review of Architecture, Mechanisms, and Role Modelling in Education with Generative AI.

Embed this page

Copy the code below to embed a chromeless version of this page in a learning management system or other website. The embedded view hides the site header, navigation, and footer.