Concept
Medical and Health Professions Education
Medical and Health Professions Education (HPE) — the teaching and training of medical, nursing, pharmacy, and allied health professionals. AI is reshaping this domain through clinical Simulation, reinforcement learning trainers, adaptive learning, and the application of foundational learning principles (experiential, situated, and distributed cognition) in health-professions contexts. Because HPE is high-stakes, competency-based, and clinically embedded, it raises distinct questions about AI's role in skill acquisition, patient safety, and the educator's judgment. Nursing, the most heavily studied program in HPE, now has a page of its own — Nursing Education — where the competence boundary, identity-formation, and workforce evidence that this page carries only in outline is developed in full.
Questions to Consider
- In medicine, AI benefits like scalable practice and adaptive feedback must be balanced against risks like erosion of hands-on clinical skill — where errors carry direct patient consequences. Where would you draw the line between what AI should do and what a trainee must practice for themselves?
- The page argues AI should be used to operationalize age-old learning principles — experiential, situated, distributed cognition — rather than replace the educator's guiding role. What makes the educator's judgment indispensable even when AI can simulate or personalize the practice?
- Reinforcement-learning trainers and agentic AI are now used for clinical and procedural skills in residency. If an AI agent trains a resident on a procedure, how would you verify they've actually learned it safely before they do it on a patient?
- Because health-professions education is high-stakes and competency-based, assessment questions carry particular weight. How might AI-assisted assessment both improve and threaten the evaluation of clinical competence?
- Over-reliance on AI is a specific concern in medicine. How do you think training with AI could produce a clinician who is more confident but less able to reason independently — and what could guard against that?
Introduction
AI in medical and health-professions education is a growing strand of the knowledge base's subject-area coverage. Unlike general higher education, HPE is oriented toward the development of clinical competencies, procedural skills, and professional judgment, which shapes how AI tools are designed and evaluated.
How AI appears in health-professions education
-
Operationalizing foundational learning principles. Fowlin et al. (developed at the Medical University of South Carolina) argue that AI should be used to operationalize age-old learning principles — Dewey's experiential learning, situated cognition, and distributed cognition — rather than replace the educator's guiding role. AI enhances personalized and adaptive learning while the teacher remains central to student engagement and outcomes.
-
Clinical skills and reinforcement learning. ResidencyRL uses reinforcement learning to train clinical reasoning and procedural skills in residency, demonstrating AI as a skills-training partner in real clinical workflows.
-
Simulation and agentic AI. Agentic AI simulation in brachytherapy shows how AI agents support hands-on procedural training in medical specialties.
-
Gamification of medical learning. MedGame applies gamification and LLMs to engage medical students.
-
Nursing and interdisciplinary education. AI transformation of nursing education documents how AI reshapes nursing curricula and instruction. This page keeps nursing as one program among medicine, pharmacy, dentistry, and allied health, organized around the shared problem of scalable clinical competence; the nursing-specific record — a competence boundary where AI's documented benefits stop, an explicitly professional-identity framing, and students' workforce anxiety — belongs to Nursing Education and is cross-referenced here rather than restated.
-
AI-powered simulation in nursing. Jiang et al. (2026)'s mixed-methods systematic review of 19 studies (N = 1,253) finds AI-driven simulations (GenAI/LLMs, virtual patients/mannequins, AI-enhanced VR/MR, and chatbots) significantly improve cognitive knowledge and affective outcomes (Self-Efficacy, communication confidence) in the strongest designs, but are inconsistent for complex psychomotor skills — one RCT found AI-assisted simulation inferior to standardized patients. Their concept of an "authenticity gap" — a learner-perceived shortfall in emotional resonance, nonverbal cues, and tactile examination — grounds why AI is best for highly structured objectives (foundational communication, history-taking) and should sit in a stepped simulation continuum alongside, not instead of, human-standardized patients and clinical placement. This is a distinctive, evidence-anchored refinement of the domain's Simulation strand and parallels the acceptance-trust finding that learners favor human feedback perceived as "benevolent" over AI seen as merely "competent."
-
LLMs and the reconstitution of professional identity in nursing (2026). Sun et al. (2026)'s critical integrative review of 489 studies across 47 countries (Whittemore & Knafl framework) shifts the question from what LLMs can do in nursing education to what their integration does to the developmental processes through which a nurse is formed. Framed through Benner's skill acquisition, cognitive load theory, automation bias, and Wenger's identity formation, they find the same technology both enhances and erodes competence depending on whether the displaced cognitive work is extraneous to, or constitutive of, the target competence — unstructured reliance produced measurable deficits in ethical reasoning and clinical judgment. Their Professional Identity Tension Model distinguishes LLM integration that supports professional formation from integration that silently substitutes for it, and Structural Empathy Suppression reframes apparent AI "outperformance" on relational metrics as a symptom of systemic overwork rather than genuine AI empathy. An evidence gap map shows the highest-policy-consequence domains (professional identity, relational/ethical competency, long-term outcomes) rest on the thinnest rigorous evidence — a caution against confident deployment claims in nursing education.
-
Multi-agent AI standardized patients in clinical interview training. In a 2026 randomized controlled trial with 95 analyzed medical students (Yang et al.), multi-agent AI standardized patient training raised final OSCE-aligned examination scores over a structured non-LLM progressive-disclosure control (71.8% vs. 55.6%; β = 16.4 percentage points; P = 3.30e-4) while leaving binary diagnostic accuracy statistically identical (84% vs. 86%; P = 1.000). The dissociation indicates that early clinical-training gains can surface first in consultation process quality — communication (mean 3.53 vs. 2.64 on a 1–5 OSCE scale; P = 4.50e-4), empathic expression (+31 percentage points on the "expressing empathy" checklist item), and selected history-taking behaviors — before any improvement in diagnostic endpoints, and that AI patients still need human raters and faculty for professionalism and readiness judgments.
-
Students default to treating an AI patient as a searchable database. Voice-interviewing a genAI avatar in a PBL pilot, students asked closed-ended, efficiency-focused questions — one asking whether the AI was "just, like, a question base" — and the encounters ran 55–65 minutes versus 36–39 for the legacy module (Mool et al., 2026).
-
Task allocation and the SCAN framework. Tsim et al. reframe AI integration from learner "misuse" to misclassification — a failure of real-time metacognitive evaluation. Their SCAN framework (Substitute, Complement, Aid, Non-Negotiable), grounded in Vygotsky's zone of proximal development, allocates generative AI tasks by the individual learner's developmental state and identifies passive engagement within AI-scaffolded tasks as a hidden pathway to mis-skilling that requires re-identification from AI to expert assistance with human epistemic auditors.
-
Human-in-the-loop instructional asset generation. Gen-Mentor (Dong et al. 2026) integrates a vision-language-model backbone into a dental-radiography workflow: Faster R-CNN localizes four target findings (Filling, Implant, Impacted Tooth, Cavity), a conditional diffusion model generates class-specific synthetic ROI candidates, a VLM produces evidence-linked captions, and an Large Language Models (LLMs) reformats them into case descriptions, comparisons, and quiz prompts — all before structured expert review. Evaluated with 45 dental students (mean SUS 72.7), it demonstrates how human-in-the-loop review of AI-generated instructional assets can expand case diversity and immediate-feedback support while retaining expert oversight over what students see.
-
AI grading of pharmacy exams. Falahat, Das, Bhaumik & Thambi (2026) evaluated ChatGPT-5 against human faculty grading of a 21-item pharmacy exam (16 students) across multiple-choice, select-all-that-apply, fill-in-the-blank, listing, short-answer, and essay items. The model matched faculty closely on objective items (CCC 0.935–1.000) but was unreliable for listing, short-answer (CCC ≈0), and essay (0.341–0.854) responses, and a structured rubric did not consistently improve agreement — evidence that in high-stakes, competency-based Assessment in health professions, AI suits well-specified items while human review remains necessary for subjective, open-ended clinical reasoning.
-
A validated LLM judge for clinical communication. Hasan et al. (2026) scored the 3E communication skills from transcripts alone with a judge that reached Pearson 0.759 and ICC(A,1) 0.746 against the consensus of human raters, inside the spread of individual humans, though its weakest dimension (Be Explicit) did not improve across encounters.
-
GenAI in scenario-based healthcare education. Neto and colleagues (2026) systematically reviewed 23 studies of GenAI across scenario-, case-, problem-, and simulation-based learning in healthcare education (PRISMA 2020). Their central finding is that prompt design functions as instructional specification — encoding the cognitive targets and quality criteria implicit in expert authoring — yet only 34.8% of studies aligned generated content with instructional frameworks and only 34.8% reported prompting in enough detail to reproduce. GPT-4 dominated implementations (44.4%), hybrid human-AI collaboration outperformed fully automated approaches, and evidence was strongest for higher-order cognitive skills but inconsistent elsewhere. This grounds scenario/simulation-based and problem-based Pedagogies and Teaching Strategies in medical education with validation- and integration-quality benchmarks.
-
AI scoring of open-ended exam questions. Olvet et al. (2026) tested whether GPT-4 could reliably score open-ended questions on pre-clerkship assessments at two US medical schools. With faculty iteratively refining scoring rubrics across three rounds of error-pattern analysis, AI–faculty inter-rater reliability reached substantial-to-almost-perfect agreement on three of four questions (weighted kappa up to 0.94) but only moderate on the holistic-rubric item (κw = 0.54); discrepancies traced to both raters (GPT-4 over-scoring multiple-answer or rubric-absent-vocabulary responses; faculty being overly generous) and occasional feedback inaccuracies keep humans in the loop. The authors argue the case for automated OEQ scoring is strengthened because ~82% of US medical schools grade pre-clerkship work pass/fail, where exact AI score agreement is not always required.
-
Evaluating AI teaching agents, not just deploying them. Zhang et al. (2026) deployed eight LLM teaching agents across four role-play paradigms (patient, student, expert, family member) in an endocrinology curriculum and scored 167 student dialogues with an 8-dimension teaching-quality rubric. Platform scores diverged sharply from rubric quality; agents differed most on knowledge dimensions and least on role enactment; adaptive difficulty calibration was a shared weakness; and strengthening an agent's empathy raised role-play quality without improving knowledge coverage. The finding that role-play paradigms can be separated (beyond the default patient-doctor script) points to designing and evaluating agents against explicit pedagogical dimensions.
-
The shape of the field's own literature. Sriram, Nichols, Ganti and Gue (2026) map the ethical-AI-in-medical-education literature itself: 1,403 Web of Science publications from 1995 to 2026, negligible for two decades, then 113 in 2023, 256 in 2024 and 500 in 2025; 39 countries clear the authorship threshold, led by the United States (533 publications) over China (189) and England (95); and the most-cited works are capability tests — ChatGPT's performance on medical licensing examinations above all — rather than governance scholarship. Their reading is that the field has grown fast without growing evenly, and that empirical evaluation of governance frameworks is the work the citation record is not rewarding. It is a useful caution for anyone treating the volume of published AI-in-medical-education research as evidence that its safe-use questions have been answered.
A 153-report scoping review of GenAI in medical education maps applications — simulation and clinical skills largest (47 reports), ahead of assessment generation and feedback (22) and case-based learning and clinical reasoning (21) — yet only 13 reports carried a follow-up or retention signal and no patient-level outcome was found (Zhao et al. (2026)).
Why it matters
HPE is a high-stakes, competency-based domain where AI's benefits (scalable practice, adaptive feedback, simulation) must be balanced against risks (Over-Reliance, erosion of hands-on clinical skill, ethical and safety concerns). The knowledge base's general concepts — Teaching, Assessment, Feedback, Equity, and Ethics — apply with particular intensity in health professions, where errors carry direct patient consequences.
Wang and Shan (2026) name the resulting "safety gap" — the divergence between a student's AI-assisted performance and their unassisted ability to verify it — and prescribe withholding solutions, constructive cognitive friction, and Socratic or Adversarial AI architectures for high-stakes clinical training.
Implications for health-professions educators
- Use AI to operationalize learning principles, not replace the educator. Fowlin et al. argue AI should operationalize experiential, situated, and distributed-cognition learning while the teacher remains central to engagement and outcomes.
- Leverage AI for clinical skills training. ResidencyRL and agentic simulation show AI as a skills-training partner in real clinical workflows — embed it where it adds safe, scalable practice.
- Balance high-stakes benefits against over-reliance. HPE is competency-based and high-stakes; guard against AI substituting for hands-on clinical skill and judgment, and apply Feedback, Assessment, and Ethics considerations with particular care.
- Adapt gamified and interdisciplinary AI thoughtfully. Gamified LLM learning and nursing-education transformation show promise but need evaluation for safety and skill outcomes; for nursing specifically, that safety-and-skill evidence — including the RCT in which AI-assisted simulation underperformed standardized patients — is gathered on Nursing Education.
Connected Concepts
- Problem-Based Learning
- Higher Education
- Simulation
- Adaptive Learning
- Personalized Learning
- Game-Based Learning
- Experiential Learning
- Situated Learning
- Distributed Cognition
- Teaching
- Assessment
- Feedback
- AI in Education
- AIEd in the Disciplines
- Nursing Education — the nursing strand of health-professions education
- Virtual and Augmented Reality — immersive and AR clinical training
Connected Articles
-
Generative artificial intelligence and the transformation of medical education: a scoping review — Scoping review of 153 medical-education GenAI reports: simulation leads, durability and patient outcomes unmeasured
-
Evaluating Scaffolding-Oriented Multi-Agent Large Language Model System for Clinical Interview Training — Evaluating Scaffolding-Oriented Multi-Agent Large Language Model System for Clinical Interview Training
-
What Platform Scores Miss: Multidimensional Evaluation of AI Teaching Agents in Medical Education — 8-dimension rubric evaluation of AI teaching agents in medical education
-
When the algorithm enters the classroom: A critical integrative review of large language models, nursing education — LLMs, nursing education structural gaps, and the reconstitution of professional identity (Sun et al. 2026)
-
AI as Teammate: Rethinking Task Distribution in Medical Training — SCAN framework: rethinking AI task distribution in medical training (Tsim et al. 2026)
-
Empowering Educators: Operationalizing Age-Old Learning Principles Using AI — Operationalizing experiential, situated, and distributed cognition with AI in health-professions education
-
ResidencyRL: Reinforcement Learning in Simulated Clinical Environments — Reinforcement-learning training for clinical skills in residency
-
MedGame: Storytelling Gamification Empowered by Large Language Models for Medical Education — Gamified LLM-based learning for medical education
-
Agentic AI-driven Immersive Simulation: A Knowledge-Aware Virtual Training Platform for High Dose Rate (HDR) Brachytherapy — Agentic AI simulation for brachytherapy training
-
Transforming Nursing Education with Artificial Intelligence: A Systematic Review (2010–2025) — Transforming nursing education with AI
-
AI-Powered Simulation for Nursing Education: Mixed Methods Systematic Review — AI-powered simulation in nursing: mixed methods systematic review
-
The Safety Gap: Restoring Productive Struggle Through Pedagogically Aligned Generative AI — The Safety Gap: Restoring Productive Struggle
-
Gen-Mentor: A Human-in-the-Loop Instructional Framework for Dental Radiography Using Generative AI — Gen-Mentor: human-in-the-loop dental radiography instruction (Dong et al. 2026)
-
Generative AI in Scenario-Based Healthcare Education: A Systematic Review of Applications, Validation Practices, and Pedagogical Integration — Systematic review of GenAI in scenario-based healthcare education (Neto et al. 2026)
-
Bridging technology and education: The use of ChatGPT in grading pharmacy student exams
-
Scalable AI-based clinical communication training and automated assessment — Scalable AI-based clinical communication training and automated assessment
-
Ethical and responsible use of artificial intelligence in medical education — Bibliometric map of AI ethics in medical education: output rose 113 (2023) to 500 (2025), 39 countries, and citations reward capability testing over governance (Sriram et al. 2026)