Markos Stamatakis, Omkar Gavali, Joshua Berger, Christian Wartena, Anett Hoppe, Ralph Ewerth (2026) โ arXiv preprint ๐ Full text (arXiv)
Investigates LLM chatbots' performance on 70 MCQs for a university CS lecture on interactive visual data analysis, comparing with student performance. GPT-4o and GPT-5 significantly outperformed smaller models. A user study in two courses showed that presenting ChatGPT answers with explanations did NOT generally improve student performance.
Key Contributions
- ChatGPT answers with explanations do not improve student MCQ performance; GPT-4o/5 outperform smaller models significantly.
Related Pages
- ai-k12-evidence-base โ Empirical evidence on AI in K-12 education outcomes
- intelligent-tutoring โ Automated tutoring systems and their evaluation
- student-experience โ Student perspectives on AI in education
- learning-analytics โ Data-driven approaches to understanding learning
- llm-feedback-programming-classroom โ LLM feedback in classroom settings
Citation
APA: Markos Stamatakis, Omkar Gavali, Joshua Berger, Christian Wartena, Anett Hoppe, Ralph Ewerth (2026). Are LLM-based Chatbots Good Enough to Support Computer Science Students in Multiple-Choice Exercises?. arXiv:2606.15919. arXiv preprint.