Research Article
From Traditional Classroom to AI-Enhanced Flipped Classroom: A Three-Year Pedagogical Evolution for International Students in Pharmacology
Synthesis: A three-year quasi-experiment at one Chinese medical university compared traditional, flipped, and AI-enhanced flipped instruction in pharmacology for 68 international clinical-medicine undergraduates across three cohorts (n = 21 in 2023, n = 16 in 2024, n = 31 in 2025). Generative AI was layered onto the flipped model as a cognitive scaffold for pre-class preparation and presentation work — retrieving, outlining, and explaining concepts — with students required to verify, rewrite, and take responsibility for every output. The AI-enhanced cohort showed the strongest classroom numbers, and unlike the standard flipped cohort it also beat the traditional classroom on assignment accuracy; yet it did not surpass the flipped cohort on in-class presentation or deep mastery of professional knowledge. The pattern positions human–AI collaboration as a process accelerator that frees instructor time for higher-order work rather than a substitute for it.
Key Findings
- 68 third-year international clinical-medicine undergraduates were followed over three years, one model per year: traditional (2023, n = 21), flipped (2024, n = 16), AI-enhanced flipped (2025, n = 31).
- Response rate was significantly higher in the AI-enhanced group than the flipped (p < 0.01) and traditional (p < 0.05) groups.
- Question accuracy beat the traditional classroom in both the AI-enhanced (p < 0.01) and flipped (p < 0.05) groups.
- Assignment accuracy was significantly higher in the AI-enhanced group than the traditional group (p < 0.01) — an advantage the standard flipped classroom missed.
- The AI-enhanced group beat the flipped group on the comprehensive assessment score and pre-class preparation and PowerPoint quality (p < 0.05), but not on in-class presentation or professional-knowledge mastery.
- The post-course questionnaire drew 29 valid responses out of 31 (93.55%); 75.86% used AI weekly, and the highest-rated item was verifying AI content against textbooks or literature (4.07, 79.31%).
- AI rated strongly for explaining mechanisms and stimulating inquiry into drug–drug interactions (3.93 each) but was trusted less for factual reliability (3.34, 48.28%); 51.72% (n = 15) named content needing substantial rework as their main challenge.
The three cohorts and the three instructional models
The setting was Kunming Medical University's English-taught Bachelor of Medicine International Stream (BMEIS) program, which has enrolled international students from dozens of countries since 2011. All three models ran across pre-class, in-class, and post-class phases, shifting from teacher-led transmission to self-directed active learning to a human intelligence–AI collaborative framework. Content, credit hours, and assessment were identical across the three years; median baseline scores were comparable at 91, 92.3, and 92 (p > 0.05), as were male proportions (47.62%, 75.00%, 48.39%) and median ages (21, 22, and 21 years). The small cohorts reflect the COVID-19 pandemic of 2019–2022.
Classroom interaction and academic performance
Response rate in the AI-enhanced group exceeded both the flipped (p < 0.01) and traditional (p < 0.05) groups, and question accuracy was higher than traditional in both the AI-enhanced (p < 0.01) and flipped (p < 0.05) groups. Its assignment accuracy was significantly higher than the traditional group's (p < 0.01), while the flipped group showed no such post-class advantage. The comprehensive assessment score (100 points: preparation and PowerPoint quality 30, in-class presentation 40, professional-knowledge mastery 30) and the preparation subscore were both significantly higher than the flipped group's (p < 0.05); in-class presentation and professional-knowledge mastery did not differ. The AI layer improved engagement and preparation, not deep outcomes.
How students used and evaluated the AI tools
The 2025 cohort used two recommended free tools — DeepSeek V3.1 (released 21 August 2025) and Kimi's OK Computer Agent (September 2025). Students could not submit AI content directly; they had to verify and integrate outputs with authoritative sources, and instructors gave only brief oral guidance, no formal training or prompt templates. The anonymous self-report questionnaire used 5-point Likert items (Cronbach's α of 0.84, 0.79, and 0.84). Efficiency dominated perceived value: saving time on information collection (mean 3.79, 75.86%) and making output more professional and concise (3.66, 68.97%), with visualization suggestions rated most conservatively (3.14, 51.73%). Intentions were favorable: 82.75% agreed on combining instructor guidance with AI, 79.31% would verify content against literature, and 75.87% intended to keep using AI.
Mechanisms and boundaries
Results are explained through Constructivism learning theory, cognitive load theory, the technology acceptance model, and self-regulated learning: AI reduced extraneous load by handling retrieval and summarization. The gains had boundaries: preparation and interaction improved without surpassing the standard flipped model on in-class presentation or deep mastery. Students' expectations agree: 75.86% (n = 22) selected all three roles of personal tutor, creative partner, and critical thinking partner, 72.41% (n = 21) recognized a research-assistant role, and only 55.17% (n = 16) saw AI primarily as a content generator. The model is judged transferable to other knowledge-intensive fields such as pathology and internal medicine, though skill-based disciplines need hands-on components and low-cost tools.
What this means for practice
- Instructors. Treat AI as a process scaffold and reserve instructor-led time for the higher-order work it does not reliably deliver.
- Instructional designers. Build verification rules into the task rather than banning tools: students had to check AI output against textbooks and literature and rewrite it themselves.
- Administrators. Budget and guidance matter more than tool choice; benefits may not replicate without low-cost tools and basic digital infrastructure.
- Course teams. Expect gains in interaction and preparation quality first, and design separate assessments if deep knowledge mastery is the target.
- Faculty developers. The absence of formal training was a research choice, not a recommended rollout; planned guidance may reduce the heterogeneity the authors flag.
Limitations
- The quasi-experimental design assigned cohorts by academic year, not randomly, so cohort effects cannot be ruled out; the study covered one pharmacology course at one institution, and students were from South and Southeast Asia, comfortable with English-medium instruction.
- The traditional (2023) and flipped (2024) groups were not surveyed about AI use, and earlier cohorts may have accessed other assistants, which could bias results toward the null.
- No structured training was provided, so heterogeneity in self-directed use patterns may have added variability.
- The AI tools were the 2025 generation — DeepSeek V3.1 (released 21 August 2025) and Kimi's OK Computer Agent (September 2025) — since superseded, so the effects reflect that 2023–2025 window.
Citation
Liu, Y. J., Ma, H. Z., Luo, H. Y., Xiong, Y. X., Qu, F. W., Xie, J. P., & Guo, Y. (2026). From traditional classroom to AI-enhanced flipped classroom: a three-year pedagogical evolution for international students in pharmacology. BMC Medical Education.