Research Article
Ethical and responsible use of artificial intelligence in medical education
Synthesis: Sriram, Nichols, Ganti and Gue map the literature on ethical and responsible use of AI in medical education with a bibliometric analysis of 1,403 Web of Science publications from 1995 to 2026, using VOSviewer to build country, institution and keyword co-occurrence networks. The field was negligible for two decades, rose gradually after 2019, and then inflected sharply after ChatGPT's release: 113 publications in 2023, 256 in 2024, 500 in 2025. The map's most telling feature is what the citation record rewards. Output is concentrated in 39 countries meeting the authorship threshold, with the United States far ahead (533 publications), then China (189) and England (95), and Harvard and its affiliates leading institutional output — while the most-cited works are capability-testing studies, above all ChatGPT's performance on medical licensing examinations. The authors' conclusion is that the field has grown fast without growing evenly, and that governance-focused scholarship has not kept pace with the studies that measure what the models can do.
Key Findings
- The search of the Web of Science Core Collection, combining medical-education and AI-related terms and limited to 1995–2026, returned 1,403 publications; networks were built with VOSviewer 1.6.20 and descriptive figures with Excel.
- Output was negligible for nearly two decades, rose gradually after 2019, and then accelerated: 17 publications in 2020, 29 in 2021, 40 in 2022, 113 in 2023, 256 in 2024 and 500 in 2025 — an inflection the authors tie to widespread generative-AI adoption after ChatGPT's public release, while noting the underlying disciplinary concern about algorithmic opacity and reproducibility predates it by several years.
- Geographic analysis identified 39 countries meeting the 10-publication inclusion threshold for the authorship network, with the United States the most influential and highly connected node on 533 publications, followed by China (189) and England (95); Canada, India, Germany, Saudi Arabia, Australia, Turkiye and Italy were also prominent.
- Harvard University and its affiliated entities led institutional output, while Harvard Medical School, Stanford and the University of California San Francisco led on citation strength.
- The highest citation impact belongs to Tjoa and Guan's 2021 survey of explainable AI in medicine (1,475 total citations, 245.83 per year), followed by Dave et al.'s 2023 article on ChatGPT's applications and ethical considerations in medicine (861 citations, 215.25 per year); most top-cited items appeared between 2021 and 2024.
- The keyword co-occurrence network (49 keywords above the 20-occurrence threshold) is led by "artificial intelligence" (n = 644), "medical education" (471), "chatgpt" (230) and "large language models" (134), and splits into three clusters: traditional AI and predictive modeling, generative-AI and LLM applications, and early generative-AI discussions in medicine.
- The authors' assessment is that citation impact is dominated by capability-testing rather than governance scholarship, and they call for empirical evaluation of AI governance frameworks plus broader institutional and geographic representation in the field.
What a citation map of this field shows
Bibliometric analysis answers a different question from a systematic review. It does not appraise study quality or synthesize findings; it reads the shape of a literature — who publishes, who is cited, which terms co-occur — and that shape is the finding here. The growth curve is the clearest signal: a field that barely existed before 2019 acquired half of its output in two years, with the sharpest step in the year after ChatGPT became publicly available. The authors are careful about causality, noting that concerns about algorithmic opacity and the "black-box" character of clinical decision support had been accumulating through the late 2010s, so generative AI gave an existing debate a more visible target rather than creating it.
The concentration is the second signal. Only 39 countries cleared the 10-publication threshold for the authorship network, and the United States accounts for more than twice China's output while the top institutions cluster in a handful of American medical schools. For a field whose subject matter is how AI should be governed in clinical training, a map drawn mainly from a few national systems is itself a limitation on what the scholarship can see — the deployment contexts where AI governance questions are hardest (resource-constrained systems, non-English clinical settings) are the least represented.
The gap between what is cited and what is needed
The most-cited items are not ethics frameworks. They are capability tests: ChatGPT's performance on medical licensing examinations and similar demonstrations of what a model can do. That is a coherent research economy — a new tool invites immediate measurement — but it produces a literature that answers "can it?" far more often than "should it, under whose oversight, and with what recourse for the student or patient affected?". The authors read the citation record as evidence that governance-focused scholarship is underdeveloped relative to capability testing, and their recommendation follows directly: evaluate governance frameworks empirically rather than proposing them, and widen the evidence base beyond the institutions that currently dominate it.
What this means for practice
- Read a growth curve as an agenda, not a validation. The field's rapid expansion tracks tool availability, not accumulated evidence about safe use. When adopting an AI tool in a medical curriculum, the volume of published work on it is not evidence that its governance questions have been answered.
- Treat governance as an empirical question. The authors' own recommendation is that AI Governance frameworks should be evaluated rather than asserted; that means asking which policy actually changed what a student or patient experienced, and recording it.
- Check who the evidence comes from. With the authorship network concentrated in a few countries and institutions, findings about AI in clinical training carry context that may not transfer — particularly for programs outside the English-language, high-resource settings that dominate the corpus.
- Pair capability studies with consequence studies. The citation record rewards measuring model performance on licensing exams. The complementary work — what changes in learning, supervision, error detection or equity when those models enter a curriculum — is what the map shows is missing.
Limitations
- The corpus comes from a single database (Web of Science Core Collection) with a keyword-based search strategy, so coverage depends on indexing and on how authors and journals phrase their work; publications indexed elsewhere, and non-English work, are systematically less visible.
- Bibliometric indicators measure attention rather than quality: a highly cited capability study and a highly cited governance framework count the same in a citation network, and citation lag means the 2025–2026 counts are still accruing.
- The analysis describes the shape of a literature, not what its studies found; the governance gap the authors identify is inferred from publication and citation patterns rather than from a content appraisal of the governance scholarship itself.
Citation
Sriram, A., Nichols, E., Ganti, L., & Gue, S. (2026). Ethical and responsible use of artificial intelligence in medical education. International Journal of Emergency Medicine, 19, 239.