Research Article
Evaluating Scaffolding-Oriented Multi-Agent Large Language Model System for Clinical Interview Training
Synthesis: This randomized controlled trial (N = 100 third-year medical students) tests MeduAI-SP, a generative-AI platform that organizes separate LLM agents — a simulated patient, a Socratic tutor, and a turn-level evaluator — around explicit Scaffolding functions for clinical interview training. Compared with a structured non-LLM progressive-disclosure control built from the same cases, the multi-agent condition produced markedly higher final examination scores (71.8% vs. 55.6%; Hedges' g = −0.81) and OSCE communication ratings (3.53 vs. 2.64 out of 5), while binary diagnostic accuracy was statistically identical (84% vs. 86%). The design deliberately withheld diagnostic answers and summative scores, an Intelligent Tutoring strategy intended to support the consultation behaviors that carry clinical reasoning without substituting for the learner's own reasoning. The authors frame the result as evidence for functional complementarity: AI agents handle repetitive role-play, checklist monitoring, and Socratic prompting, while human educators retain judgment, emotional interpretation, and credentialing. The study is limited by an AI-only comparator, a single acute-abdominal case family, and no delayed transfer or patient-reported outcomes.
Context and Research Gap
Clinical interviewing competence couples cognitive work — gathering structured information, generating and revising diagnostic hypotheses under uncertainty — with communicative work, including empathic expression that shapes rapport, disclosure, and the completeness of elicited information. Case-based learning, computerized virtual patients, and simulated-patient encounters each address part of this, but traditional standardized-patient training resists scale because it requires trained actors, faculty supervision, physical space, fixed scheduling, and substantial institutional resources. The stated gap is not whether LLM agents can imitate patients — fidelity studies are accumulating — but whether organizing agents around specific instructional functions improves the process quality of simulated interviews relative to structured non-LLM materials derived from the same cases. The authors position the work against two risks of Generative AI in clinical training: over-reliance that weakens critical thinking, and AI-induced "never-skilling," in which trainees who lean on AI too early fail to build the independent reasoning safe practice requires. They ground the design in Scaffolding theory and the zone of proximal development, and in cognitive load theory's argument that support should be timely, structured, and phased without displacing the learner's own reasoning.
System Design: Scaffolding-Oriented Multi-Agent Architecture
MeduAI-SP (Medical Education Artificial Intelligence–Standardized Patient) is a web-based platform (React/TypeScript frontend, FastAPI backend, PostgreSQL, powered by the qwen-max model) that exposes four modes: single-agent learning, multi-agent learning, structured control learning, and examination. The multi-agent condition pairs four roles. A patient agent generates inquiry-dependent responses from a structured YAML case script and is instructed to avoid premature disclosure of diagnostic clues. A teaching (tutor) agent provides Socratic prompts covering four areas — focused history taking, diagnostic reasoning, summarizing/confirming information, and rapport-building communication — and is explicitly barred from revealing the final diagnosis, key differentials, or case answers; prompts are phrased as reflective questions such as asking whether the current information suffices to support a conclusion, or reminding the student to acknowledge distress before continuing. A turn-level evaluator monitors coverage of key history-taking and communication items and detects when a learner is stuck, missed critical information, impaired rapport, or made an unsafe inference; the tutor fires only on a flagged need such as missing key history, conversational impasse, premature closure, communication breakdown, or professionalism concerns. A final evaluator composes a post-encounter OSCE-aligned summary that stays invisible during the live interaction, keeping assessment formative rather than a revealed score. This is an Agentic AI architecture in which role specialization, not raw model capability, carries the pedagogical load.
Study Method: Randomized Controlled Trial
The design was a two-arm parallel randomized controlled trial. One hundred volunteer third-year clinical-medicine undergraduates at Guangzhou Medical University (age 20–26, mean 22.3; 63 women, 37 men) were randomized 1:1 by computer-generated sequence; 95 (MA n = 47; CT n = 48) completed the full workflow and entered the complete-case primary analysis. Both arms ran two learning encounters, one examination encounter, and a post-intervention survey on the same schedule, cases, and assessment method. The comparator was deliberately non-LLM: a computer-based progressive-disclosure activity using the same authored case materials in fixed stages with predetermined answers, removing AI dialogue, tutor scaffolding, real-time process monitoring, and automated feedback. The examination was identical for both arms — a patient-only AI standardized patient with teaching and process-monitoring agents hidden — and no real-time instructional feedback was given during it. Three acute abdominal cases were used: acute appendicitis and acute pancreatitis as learning cases, and perforated peptic ulcer as the examination case, drawn from a platform library of 18 standardized cases. The primary outcome was final OSCE-aligned examination performance, combining an overall score, four global 1–5 domain ratings (history taking, clinical thinking, communication and empathy, diagnostic accuracy), and a weighted behavior checklist; secondary outcomes covered domain scores, checklist coverage, diagnostic accuracy, worksheet completion, and usability/engagement via the System Usability Scale (SUS) and User Engagement Scale (UES).
Results: Process Quality Improved, Diagnostic Accuracy Unchanged
The multi-agent condition raised final examination performance: mean 71.8% vs. 55.6% (medians 78.8% vs. 60.6%), Mann–Whitney P = 5.51 × 10⁻⁵, Hedges' g = −0.81 (the negative sign reflects comparison coding). In a linear model the effect was β = 16.4 percentage points (95% CI 7.7 to 25.1; P = 3.30 × 10⁻⁴), with a smaller advantage in weighted checklist score (β = 11.7; 95% CI 0.2 to 23.1; P = 0.046).
The largest domain-level gain was in communication: 3.53 vs. 2.64 on the 1–5 scale (P = 4.14 × 10⁻⁴; g = −0.79; regression β = 0.90; 95% CI 0.41 to 1.39; P = 4.50 × 10⁻⁴). History-taking completeness trended positive but missed significance (β = 0.34; 95% CI −0.05 to 0.73; P = 0.091), and clinical reasoning differences were small. At item level the single largest gap was the checklist behavior "expressing empathy," 31 percentage points higher in the multi-agent group and still significant after Holm correction (P = 8.30 × 10⁻⁴); medication-allergy history survived Holm correction, while duration of illness and summarizing/confirming information reached significance under FDR correction.
By contrast, binary diagnostic accuracy did not differ — 86% correct initial diagnoses in control vs. 84% in multi-agent (P = 1.000) — and worksheet completion was likewise comparable (P = 0.729). The authors read this as evidence that the experimental control worked: since both arms saw the same diagnostic content, the intervention's benefit came from how learners conducted the consultation, not from unequal disease knowledge. Exploratory UMAP + k-means (k = 5; silhouette 0.34) identified five examination performance phenotypes whose distribution differed across arms (χ² = 16.15; P = 0.040), suggesting the intervention shifted students toward different profiles rather than lifting everyone uniformly. Within the multi-agent arm, learning-session gains were not uniformly positive: weighted checklist change from session 1 to session 2 averaged −14.4 percentage points (median −4.5) across 46 participants, hinting at heterogeneous trajectories. Median messages per session were 20.0 (learning 1), 30.0 (learning 2), and 17.0 (examination). Usability and engagement were favorable and broadly similar — SUS 66.9 (MA) vs. 63.2 (CT); total UES 69.8 (MA) vs. 66.9 (CT) — with the multi-agent platform slightly ahead on focused attention; survey completion was about 78% (SUS) and 58% (UES).
Corpus, Fidelity, and Scaffolding Dynamics
The released annotated corpus comprised 207 consultation sessions (118 multi-agent learning; 89 examination) and 4,815 messages: 2,350 student, 2,350 AI-SP, and 115 tutor. Median messages per session was 18 and median student utterances 9. Student dialogue was dominated by history taking, with 62.2% of utterances targeting the history of present illness and 19.1% past/family/personal history; empathic or supportive communication was only 4.3%, management/treatment discussion 5.4%, diagnostic statements 2.6%, physical examination 2.6%, and ancillary testing 0.7% — a picture of novices who concentrate on symptom gathering while underusing other parts of a complete consultation. Experts judged that 24.1% of student utterances warranted tutor scaffolding, and the need was not evenly spread: it rose from about 16.1% early to 34.8% late, suggesting support is most needed as the encounter moves toward integration, diagnostic reasoning, and management. The dominant intervention reasons were conversational impasse, domain knowledge gaps, communication breakdown, and premature closure. The AI patient itself was stable: only ~0.68% of AI-SP responses were judged to contain clear fidelity problems, and progressive disclosure was rated appropriate in ~99.3% of patient messages. Annotation reliability varied by dimension — Fleiss' κ = 0.84 for student dialogue intent (near-perfect), ICC(2,k) 0.80–0.85 across the four OSCE domains, κ = 0.56 for need-for-scaffolding, and κ = 0.31 for progressive disclosure — a pattern in which directly observable, well-defined labels agreed strongly while highly imbalanced labels did not.
Design Implications and the Boundaries of AI Substitution
The authors advance "functional complementarity" as the organizing model: AI agents are well suited to repetitive practice, consistent role-play, checklist monitoring, Socratic prompting, and preliminary formative feedback, while faculty and human standardized patients contribute contextual interpretation, nuanced emotional response, individualized remediation, professionalism assessment, and readiness judgments. Partial substitution may be acceptable for bounded tasks — an AI standardized patient absorbing some repetitive early-stage interview practice, an evaluator agent flagging checklist gaps — but the system should not autonomously decide clinical competence, make progression or credentialing decisions, or replace human-led remediation in emotionally or professionally complex cases. The boundary depends on the consequences of error, task uncertainty, relational sensitivity, and the availability of meaningful human review. A further implication for evaluation is methodological: judging AI simulated patients only by diagnostic accuracy would miss the benefits this study detected, because early gains appear first in the behaviors that make reasoning possible — asking clinically useful questions, confirming information, avoiding premature closure, and maintaining patient-centered communication. Process improvements should nevertheless be weighed alongside explicit safeguards against over-reliance and never-skilling. The authors also caution that expert-rated empathic expression is not the same as patient-perceived empathy, and cite conflicting evidence in which patients rated chatbot cancer responses more empathic than physician responses in one study yet favored physician responses in another, with perceived authorship itself shifting ratings.
Limits and Open Questions
Several limitations bound the claims. First, the comparator was a structured non-LLM progressive-disclosure activity, not a human standardized patient or faculty-feedback arm, so the results support multi-agent scaffolding over structured case materials but cannot establish that AI-SP feedback matches or exceeds human feedback — especially on empathic communication. Second, the study occupies a narrow clinical domain: cases were acute abdominal/gastrointestinal presentations (appendicitis and pancreatitis for learning; perforated peptic ulcer for examination), and the primary analysis centered on a small number of authored cases, so generalization to other specialties or authentic clinical environments should be cautious. Third, several analyses were exploratory and hypothesis-generating rather than confirmatory, including the phenotype clustering, the within-arm learning-trajectory analysis, and correlations between process variables and survey scores. Fourth, no real patients or human standardized patients participated, so whether observed communication behavior would make patients feel heard, understood, respected, or supported remains unknown. Fifth, there was no delayed follow-up, leaving open whether gains persist or transfer to higher-stakes settings, and whether Scaffolding should be faded to protect independent interviewing and reasoning from prompt dependence. These gaps point toward future comparisons of human-SP plus faculty feedback, AI-SP plus faculty oversight, and hybrid models using shared outcomes, plus validation with patient-reported communication measures and demographically diverse cohorts.
Connected Concepts
- Medical Education
- Scaffolding
- Generative AI
- Agentic AI
- Simulation
- Socratic Method
- Formative Assessment
- Intelligent Tutoring
- Human AI Collaboration
- Assessment Validity
- Human In The Loop AI
- Pedagogical LLM Training
- RCT
- Simulating Students
- Feedback
Connected Articles
- Medeasy AI Standardized Patients — MedEasy: Designing AI Standardized Patients for Clinical Consultation Training
- GenAI Simulate Patient History Pbl 2026 — Using Generative AI to Simulate Patient History-Taking in a Problem-Based Learning Tutorial: A Mixed-Methods Study
- Zhang Platform Scores Miss AI Teaching Agents 2026 — What Platform Scores Miss: Multidimensional Evaluation of AI Teaching Agents in Medical Education
- AI Teammate Task Distribution Medical Training 2026 — AI as Teammate: Rethinking Task Distribution in Medical Training
- Multi Agent Instructional Design — Multi-Agent Systems for Instructional Design
- Medgame LLM Medical Education Gamification — MedGame: Storytelling Gamification Empowered by Large Language Models for Medical Education
- Astra Multi Agent Tutoring Benchmark 2026 — ASTRA: A synthetic benchmark for trace-based evaluation of socially intelligent multi-agent tutoring
- GenAI Scenario Based Healthcare Education 2026 — Generative AI in Scenario-Based Healthcare Education: A Systematic Review of Applications, Validation Practices, and Pedagogical Integration
- Llms Misconception Collaborative Learning Healthcare 2026 — Implementing LLMs to Support Misconception-Based Collaborative Learning in Health Care Education
- AI Use Critical Thinking Medical Students 2026 — From AI Use to Critical Thinking Among Medical Students: A Moderated Mediation Perspective on Cognitive Load and Self-Regulated Learning
Citation
Yang, L., Liu, H., Li, S., Jia, R., Xiao, Y., Chen, G., & Lu, L. (2026). Evaluating Scaffolding-Oriented Multi-Agent Large Language Model System for Clinical Interview Training. arXiv preprint arXiv:2609.10939.