Research Article
Bridging technology and education: The use of ChatGPT in grading pharmacy student exams
Bridging technology and education: The use of ChatGPT in grading pharmacy student exams — This mixed-format study evaluates ChatGPT-5 against human faculty (gold standard) grading of a 21-item pharmacy exam completed by 16 students across multiple-choice, select-all-that-apply, fill-in-the-blank, listing, short-answer, and essay questions, testing two rubric conditions and two submission formats. It finds near-perfect Automated Assessment agreement for objective items but unreliable agreement for subjective open-ended responses, and shows that providing a rubric did not consistently improve LLM grading. The work advances understanding of when generative AI can substitute for human grading in Higher Ed assessment versus where human oversight remains necessary.
Key Findings
- ChatGPT-5 achieved substantial to near-perfect concordance with faculty grading for objective item types (multiple-choice, select-all-that-apply, fill-in-the-blank; CCC = 0.935–1.000), regardless of rubric use — though correct answers were provided during grading, which likely contributed to this performance.
- Accuracy and concordance declined markedly for listing (CCC 0.621–0.708), short-answer (CCC ≈ 0 to negative), and essay questions (CCC 0.341–0.854), where subjective partial-credit scoring requires contextual interpretation.
- Providing a structured rubric did not consistently improve overall accuracy or agreement in full-exam grading (71.1% without vs 68.2% with a rubric; CCC 0.740 vs 0.710), differing from prior work that found rubrics help LLM short-answer scoring.
- When responses were grouped and graded by question type, rubric use improved listing accuracy and concordance (CCC 0.773 vs 0.652) but not short-answer or essay items, where rubric-free grading often showed higher agreement — suggesting AI grading performance is sensitive to grading context and rubric design.
- The study highlights a methodological distinction between scoring accuracy and agreement: moderate percent accuracy frequently coexisted with low concordance correlation coefficients, limiting AI's reliability as a grading substitute.
- Authors conclude AI is strongest for objective or highly structured items, with human review remaining important for complex, subjective, or high-stakes assessments, and recommend future hybrid grading approaches.
Connected Concepts
- Automated Assessment
- Medical Education
- LLM
- Generative AI
- Higher Ed
- Automated Essay Scoring
- Assessment Validity
- Human In The Loop AI
Connected Articles
- Automated Formative Assessments A Level Sciences
- Ground Truth Reliability AIED
- LLM Formative Feedback Systematic Review 2026
Citation
Falahat, S., Das, J., Bhaumik, D., & Thambi, M. (2026). Bridging technology and education: The use of ChatGPT in grading pharmacy student exams. Currents in Pharmacy Teaching and Learning, 18, 102707.