Research Article
Learning with pedagogical agents in extended reality: A conceptual research-based framework of agent-centered learning in immersive media environments (ACLIME)
Synthesis: Ross and Kaspar propose ACLIME, a conceptual framework for agent-centered learning in immersive media: Pedagogical Agents embedded in Virtual and Augmented Reality as the focal point of learner interaction. Three layers: a technical foundation of high immersion, multimodal body-based interactivity, and the realism of agent and learner representation; instructional design principles moderating how those possibilities become design features; and four psychological clusters — the learner's virtual self-experience, the experience of the agent (social presence, human-agent space, agent awareness), social interaction with the agent, and immediate learning factors. Unlike CAMIL, CATLM-VR and TICOL it keeps the agent in the model, covers AR as well as VR, and integrates Learning Theories as temporal dynamics. It is a synthesis, not an empirical test, leaving out long-term outcomes and learner characteristics.
Key Findings
- The framework names its own parts, and its clusters carry named dimensions. Immersion, body-based interactivity and agent and learner realism feed three psychological clusters plus immediate learning factors; design principles moderate, and outcomes such as transfer are factored out. Virtual self-experience: physical presence, self-presence, embodiment, cybersickness. Experience of the agent: social presence, human-agent space (replacing social space in Collaborative Learning), agent awareness. Social interaction: similarity to human-human interaction, task-related versus social-relational direction, intensity, persistence, the proteus effect.
- Two interaction modes define the state of the art. A tutor (guidance, reflective questioning, explanations) or a role-playing partner holding a defined role for realistic, Situated Learning-style dialogue: a climate-change field trip, a corporate negotiation Simulation.
- Immersive media differ technically in three ways. Immersion (head-mounted displays displace reality along the reality-virtuality continuum from AR through AV to VR), body-based multimodal interactivity, and representation virtualizable along the avatarization continuum.
- Realism cuts both ways. Behavioral realism is the claimed decisive predictor of social presence; agents are computer-controlled where avatars are person-controlled (virtual human agency), so low perceived agency needs high behavioral realism, and the uncanny valley threatens a sound human-agent space.
- Cognitive load is the explicit trade-off. Immersion and agent realism raise extraneous load, while agent interaction can lower intrinsic load via the collective working-memory effect and cut transaction costs; the net direction is unsettled.
- Temporal dynamics get their own layer. Presence and enjoyment rise with duration; the human-agent space matures through the four Social Penetration Theory stages; novelty may later depress motivation; the trajectory peaks early, then equilibrates above non-immersive baselines.
- Potentials and risks sit side by side (Table 1), and the research agenda is largely unfilled. Presence and embodiment are potentials against cybersickness; strong social presence helps social influence but undermines uninhibited exploration and honest self-disclosure; motivation, self-efficacy and cognitive activation against novelty-driven decline and overload. Hardly any empirical work exists, so it is a foundation for studies of mediation and moderation, longitudinal and within-session measurement, design principles, multi-agent groupings, and the outcomes it excludes.
The ACLIME framework and its components
Human-agent interaction in immersive media originates in technical features — immersion, body-based interactivity, and the realism of agent and learner representation — while environmental settings, appearance attributes and learner characteristics such as personality and prior knowledge stay outside the model.
Instructional design principles are the moderators between that foundation and the psychological dimensions: prescriptive statements abstracted into meta-requirements and instantiated as design features, since potentials are realizable and risks mitigable only through design, and such principles, being context-dependent, stay conceptual. The clusters carry the substance: physical presence, self-presence, Embodied Learning to a body ownership illusion, cybersickness; social presence, human-agent space, agent awareness; similarity to human-human interaction, direction, intensity, persistence, the proteus effect; socio-emotional reactions, Motivation, Self-Efficacy, cognitive load, cognitive activation.
The typology of pedagogical agents in extended reality
Two interaction modes come from the literature: the tutor, supporting Learners through guidance, reflective questioning and explanations, and the role-playing partner, taking a defined role for realistic, situated dialogue — a local in a climate-change field trip, a negotiating counterpart in corporate training where Feedback improves communication.
Agent characteristics are expressed through body and behavior: appearance and physical features on one side, AI-controlled or rule-based behavior on the other, with large language models enabling more flexible verbal communication; output is speech or visible text, plus gestures, gaze, posture and proxemics (Hall). Realism is visual or behavioral, and the behavioral form must carry the interaction when perceived virtual human agency is low, though more realism is not always better: the uncanny valley can cause sudden rejection.
Theoretical grounding: presence, embodiment and social agency
The spine distinguishes physical presence (spatial-perceptual self-localization), self-presence (virtual identity) and social presence (being with an intentional being whose mental states seem accessible); physical presence follows from immersion and interactivity and is expected in both VR and AR, though VR relocates the self perceptually while AR leaves the user in the real room. Embodiment is body ownership, virtual own body agency and self-location (Kilteni and colleagues): one's own real body in AR is the gold standard, the body ownership illusion is unique to immersive media, and self-presence is not equated with embodiment.
Social presence strengthens with more high-quality social cues across modalities (Oh and colleagues; Feine and colleagues' cue taxonomy), which XR renders stereoscopically and extends with haptics. Alongside sit the computers-are-social-actors paradigm, Blascovich's threshold model, and Social Penetration Theory, whose stages show self-disclosure driving the maturation of the human-agent space, where trust marks a sound state and Privacy-relevant costs slow it.
Design and research implications
The technical foundation is not a benefit in itself: Human AI Collaboration with an agent is a cognitive trade-off whose educational value depends on design choices that mitigate extraneous load while offloading intrinsic load, turning immersion, interactivity and realism into design features such as instructional strategies and feedback formats. Personality screening before certain scenarios and rapid exit strategies plus breaks between sessions follow, given after-effects lasting from around ten minutes to several hours. Agents should counteract the novelty-driven decline by satisfying relatedness, and the deep learning hypothesis extends to them.
The research gaps are mediation and moderation among agent realism, social presence, human-agent space, agent awareness and social-relational behavior; boundary conditions for the competing load effects; within-session measurement; and individual characteristics linked to outcomes.
Potentials, risks and the limits of a conceptual framework
This is a conceptual framework, not an empirical study: hardly any empirical work covers pedagogical agents in immersive media, and some relationships are extrapolated from adjacent areas — immersive interaction with avatars, or Pedagogical Agents in non-immersive media. It is deliberately partial: learner characteristics, environmental settings and Multimodal AI appearance attributes stay outside its scope, and long-term outcomes are factored out. Design principles stay conceptual, so the layer meant to convert potentials into practice is the least specified.
Several claims are unresolved, not merely untested: whether extraneous cognitive load rises or falls overall, whether social presence is moderator or mediator, and whether agent awareness raises or lowers load. Its temporal layer is admittedly hypothetical, the Figure 3 trajectory being exemplary and dependent on session duration, frequency and prior XR experience.
What this means for practice
- Designers. Cap exposure and build in a visible exit, and spend the realism budget on behavior: roughly 42% of VR headset users report cybersickness symptoms, after-effects run from around ten minutes to several hours, and higher visual fidelity risks the uncanny valley.
- Instructors. Decide per scenario whether you want pronounced social presence: it helps social influence and relationship building but works against uninhibited exploration or honest self-disclosure.
- Instructors. Screen learners for cybersickness susceptibility and keep AR as the fallback, since AR leaves the learner in the real room.
- Researchers. Instrument all four psychological clusters, because ACLIME factors long-term outcomes out and those clusters are all you can measure.
Limitations
- A conceptual review with no data of its own: hardly any empirical studies cover the field, so every relationship is synthesized, not tested.
- Part is extrapolated from adjacent research areas, and the borrowed quantitative anchors — the 42% cybersickness prevalence rate, the timing of after-effects — come from immersive-media research with no pedagogical agents.
- Scope covers one learner with one agent, excluding learner characteristics such as personality, intelligence and prior knowledge, along with environmental settings and appearance attributes.
Connected Concepts
- Pedagogical Agent — the simulated virtual character that socially interacts with a learner to facilitate learning, and the framework's focal point
- Virtual and Augmented Reality — the AR, AV and VR media whose immersion, interactivity and realism form ACLIME's technical foundation
- Embodied Learning — body ownership, virtual own body agency and self-location, the sense of embodiment that XR intensifies and the proteus effect draws on
- Situated Learning — realistic situatedness inside virtual scenarios, which the authors argue promotes the task-related direction of interaction
- Human AI Collaboration — learner and agent treated as a joint information-processing system that can share element interactivity
- Conversational AI — natural dialogue as the agent's primary verbal channel, made more flexible by large language models
- Multimodal AI — verbal, visual, auditory and invisible social cues, plus haptic channels unavailable in non-immersive media
- Simulation — role-play scenarios such as climate-change field trips, negotiations and patient consultations that the agent inhabits
- Self-Determination Theory — autonomy via virtual own body agency and relatedness via the human-agent space as routes to motivation
- Self-Efficacy — mastery experiences and realistic social feedback that immersive agent interaction is expected to strengthen
- Cognitive Psychology — working-memory limits and the three types of cognitive load that frame the framework's central trade-off
- Trust — trust toward the agent as one marker of a sound human-agent space, lowering perceived disclosure risk
- Theory Development in AI in Education — the paper's contribution as a conceptual framework intended to situate and guide future empirical work
Connected Articles
- Face value: How avatar identity shapes epistemic trust in AI-mediated learning — avatar identity and epistemic trust in AI-mediated learning
- Design and Implementation of a Real-time Multi-site Immersive Learning System Using Photon Fusion — a real-time multi-site immersive learning system, the infrastructure side of shared virtual environments
- Generative AI and Extended Reality in Collaborative Architectural Design Education: An Exploratory Studio Study — generative AI and extended reality in a collaborative architectural design studio
- From Prompt to Embodied Simulation: Using Generative AI to Create AR Physics Learning Tools — building embodied AR physics learning tools from prompts
- LumiNote: LLM-Assisted Multimodal Instruction for VR Stage Lighting Education — LLM-assisted multimodal instruction delivered inside VR
- Toward Accessible Psychotherapy Training Using AI-Driven Interactive Patient Avatars — AI-driven interactive patient avatars for clinical communication training
- Visualizing Engineering Fundamentals: Design of Mixed Reality and Physical Toolkits for Effective Learning — mixed reality and physical toolkits for engineering fundamentals
- Agentic AI-driven Immersive Simulation: A Knowledge-Aware Virtual Training Platform for High Dose Rate (HDR) Brachytherapy — an agentic immersive simulation platform for clinical procedure training
- Embodied Inquiry with AI as Facilitator: An Exploratory Case Study — embodied inquiry with AI as facilitator in physics learning
- Designing Conversational Agents for Adaptive Instructional Support in Business Simulation Gaming — conversational agents providing adaptive instructional support in business simulation games
Citation
Ross, M., & Kaspar, K. (2026). Learning with pedagogical agents in extended reality: A conceptual research-based framework of agent-centered learning in immersive media environments (ACLIME). PsyArXiv Preprints.