Research Article
What Platform Scores Miss: Multidimensional Evaluation of AI Teaching Agents in Medical Education
Eight AI teaching agents covering an endocrinology curriculum were deployed across four role-play paradigms (patient, student, expert, family) on a commercial platform. Twenty-two medical students generated 167 dialogues scored both by the platform's undisclosed algorithm and by an independently applied, expert-validated 8-dimension teaching-quality rubric (100 points). Platform and rubric rankings diverged for most agents — the agent ranked third by the platform ranked last on rubric quality, and the fourth-ranked rose to first — indicating platform scores index student performance, not agent teaching quality. Agents differed most on knowledge dimensions and shared a weakness in adaptive difficulty calibration; an empathic agent attained high role-play quality yet the lowest knowledge coverage.
Relevance to AI in Education: This is a direct demonstration that the metric an educator happens to have at hand can mislead which AI teaching agents get adopted or refined. It contributes a transparent, reusable evaluation framework and cautions against treating platform-generated scores as proxies for teaching quality, connecting to Assessment Validity and LLM-as-evaluator concerns.
Key Findings
- Platform scores diverge from rubric-measured teaching quality. Agent rankings reversed for most agents (third-by-platform ranked last on rubric quality; fourth-ranked rose to first). The two capture different constructs — student performance during the interaction versus the agent's own teaching behavior — and are complementary, not interchangeable.
- An 8-dimension rubric as a transparent external standard. Dimensions weighted by teaching function: medical knowledge accuracy and pedagogical guidance (20 pts each), knowledge coverage and role-play quality (15), adaptive difficulty and medical safety (10), student engagement and feedback quality (5). Validated against multiple LLM evaluators and a blinded medical-education expert (total-score ICC = 0.51; r = 0.61).
- Agents differ mainly in knowledge, not role performance. Between-agent variation was largest on knowledge coverage (CV=27.3%) and knowledge accuracy (18.5%), and smallest on role-play quality (5.7%). Adaptive difficulty calibration was a shared weakness across agents.
- Empathy did not ensure knowledge coverage. An agent deliberately revised to strengthen an empathic family persona attained among the highest role-play quality yet the lowest knowledge coverage and total score — affective and cognitive teaching functions do not move together.
- LLM-as-evaluator requires calibration. Three LLMs (Claude, DeepSeek, Qwen) agreed on agent ranking (rho=0.64-0.71) but differed markedly in leniency and discrimination; some were too lenient to discriminate (Qwen, DeepSeek ceiling effects). Automated scoring aligned with expert judgment on cognitive-process dimensions but poorly on subjective and affective ones.
- No detectable gender effect (male vs female students), though the analysis was underpowered and should not be read as evidence of equitable delivery.
Connected Concepts
- Medical Education
- Pedagogical Agent
- Intelligent Tutoring
- AI Ed Evaluation
- Assessment
- Simulation
- Generative AI
- LLM
- Assessment Validity
Connected Articles
- Jiang AI Powered Simulation Nursing Education 2026 — AI simulation in nursing education
- LLM Detecting LLM Generated Content Education — LLMs evaluating generated content
- GenAI Scenario Based Healthcare Education 2026 — GenAI in scenario-based healthcare education
- Pedagogy AI Mistakes — Pedagogical quality of AI output
- AI Learning Tools Engineering Education Needs — Needs- and attention-aware AI learning tools
Citation
Zhang, H., Qu, L., Zheng, J., Xiong, Y., Bai, H., Ji, R., Liu, G., Chen, W., Cheng, Z., Chen, Y., & Yang, C. (2026). What Platform Scores Miss: Multidimensional Evaluation of AI Teaching Agents in Medical Education. JMIR Medical Education, 12, e96819.