On this page

Synthesis: Sun et al. (2026) argue that large language model (LLM) research in nursing education has asked what LLMs can do while neglecting what their integration does to the processes through which a novice becomes an expert nurse. This critical integrative review of 489 studies across 47 countries reframes the evidence through an integrated lens (Benner's skill acquisition, cognitive load theory, automation bias, Wenger's identity formation), finding the same technology both enhances and erodes nursing competence depending on whether it displaces cognitive work that is extraneous to, or constitutive of, the competence being developed. The authors' Evidence Gap Map shows the domains of greatest policy consequence (professional identity, relational and ethical competency, long-term outcomes) rest on the least rigorous evidence, and they introduce the Professional Identity Tension Model and Structural Empathy Suppression to guide future research and design.

Key Findings

  1. LLMs address real structural gaps, but benefits are consistently moderated by context. LLMs help fill deficits in individualised learning support at scale, uneven access to clinical Simulation, and weak Scaffolding for progressive professional competency. Yet gains depend on the implementation framework (AI that preserves student cognitive agency, e.g. a nurse–AI collaboration RCT beating nurse-only care on patient satisfaction), on economic equity (paid-tier access barriers), and on model accuracy (RAG-customised systems lifted nursing-specific accuracy from 43.27% to 73.29%).
  2. Competency enhancement vs. competency attrition. Displacing extraneous cognitive load (documentation, routine retrieval) frees working memory for intrinsic processing; displacing intrinsic load (a full reasoning chain, an ethical justification, an individualised care plan) removes the object of learning itself. Unstructured AI reliance was linked to measurable deficits: students using ChatGPT as a sole resource scored significantly below textbook controls on ethical standards and clinical reasoning, and an AI-integrated curriculum produced higher scores yet less individualised, weaker-logic care plans — a performance–learning dissociation.
  3. The evidence base is inverted. The Evidence Gap Map found no randomised or quasi-experimental studies of professional identity or career development, only three quasi-experimental studies of relational and ethical competency, and no study with follow-up beyond 12 months, while controlled evidence concentrates in cognitive and technical outcomes. Opinion/conceptual evidence dominated the highest-stakes domains, and over 50% of studies came from just three countries.

The Professional Identity Tension Model

The review's central conceptual contribution formalises a single criterion — whether LLM-displaced cognitive and participatory work is extraneous to or constitutive of the competence being developed — across three levels: the task layer (which cognitive activities are displaced and whether they served a developmental function), the competency layer (enhancement when extraneous demands are displaced with cognitive agency preserved, vs. attrition when constitutive reasoning is displaced), and the identity layer (whether LLMs substitute for the participatory experiences through which professional identity forms). The design principle that follows: integration that offloads extraneous demand while preserving constructive or interactive engagement supports professional development; integration that displaces work constitutive of competence silently substitutes for it.

Structural Empathy Suppression

The authors propose this mechanism to reframe findings where AI appears to outperform nurses on relational metrics (e.g. an AI scoring higher than nurses on empathy ratings in a high-volume, 54.9 cases/hour context). Distinct from compassion fatigue (individual exhaustion) and emotional labour (managing emotional expression), it operates structurally: when systemic overwork makes authentic human empathic expression unsustainable and algorithmically consistent AI responses fill the gap, the AI appears comparatively empathetic — and the identity-constituting significance of nursing practice transfers to the algorithm. It reframes the policy question from "how effective is this technology?" to "what conditions are making this technology appear necessary?"

Implications for Educators and Curriculum Designers

  • Design for cognitive agency, not efficiency. Position LLMs as reasoning partners rather than answer sources; evaluate applications by their effect on students' independent reasoning, judgement, and reflection, not primarily on workload reduction.
  • Apply the dual-pathway test. LLMs are safest for displacing extraneous or routine work (documentation, information retrieval) and riskiest when they supply complete reasoning chains, ethical justifications, or individualised care plans — the cognitive work constitutive of competence.
  • Assess capability, not just performance. Outcomes measured while an LLM is available may index fluent performance rather than durable capability; use delayed, no-tool post-tests and transfer tasks on unfamiliar clinical presentations.
  • Address equity and Regulation. Mandating LLM integration without addressing access disparities (subscription costs, detection-tool bias against non-native speakers) risks widening educational inequities; frameworks should require evidence that implementations preserve developmental processes.

Connected Concepts

Connected Articles

Citation

Sun, Y., Li, H., Tao, X., Zhou, X., Gururajan, R., & Zhang, J. (2026). When the algorithm enters the classroom: A critical integrative review of large language models, nursing education structural gaps, and the reconstitution of professional identity. Computers and Education: Artificial Intelligence, 11, 100675.

Embed this page

Copy the code below to embed a chromeless version of this page in a learning management system or other website. The embedded view hides the site header, navigation, and footer.