🧠 AI Ed Wiki

Investigates LLM chatbots' performance on 70 MCQs for a university CS lecture on interactive visual data analysis, comparing with student performance. GPT-4o and GPT-5 significantly outperformed smaller models. A user study in two courses showed that presenting ChatGPT answers with explanations did NOT generally improve student performance.

Key Findings

  • The authors developed 70 multiple-choice questions (MCQs) for a university lecture on interactive visual data analysis and evaluated several LLM-based chatbots using different prompt designs.
  • GPT-4o and GPT-5 achieved the best results, significantly outperforming smaller models on the MCQ set.
  • Chatbot performance was compared with students' performance on the same questions, situating model accuracy relative to learner capability.
  • A user study in two lectures (interactive visual data analysis and computer vision) investigated how chatbot-generated answers and explanations affect students' performance.
  • The user study found that presenting ChatGPT answers together with an explanation does not improve students' performance in general — a counterintuitive result for chatbot-assisted learning.
  • Study Design & Method

    The evaluation proceeded in two phases. First, a technical benchmark: multiple LLM-based chatbots solved the 70 MCQs under different prompting strategies, with results compared against students' own performance to calibrate what "good enough" means. Second, an educational user study: students in two university CS courses were given chatbot answers with explanations, and their performance was measured against conditions without such support. This two-part design separates raw model competence from actual learning impact.

    Implications for AI in Education

    The headline implication is that model accuracy does not translate automatically into student learning: even when chatbots answer correctly, exposing students to answers plus explanations failed to improve their MCQ performance. For CS education, this cautions against treating chatbot outputs as ready-made study aids; the value of LLM support likely depends on how it is integrated into exercises and feedback. The large gap between frontier and smaller models also matters for tool selection in Higher Ed and CS Education contexts, as does the finding that MCQ-style support may need to be redesigned to produce measurable gains.

    Connected Concepts

  • CS Education
  • Higher Ed
  • Administrator
  • Pedagogical Agent
  • Automated Question Generation
  • Affective Computing
  • Agentic AI
  • AI Tutoring
  • Connected Articles

  • Cross Dataset Bloom Question Classification — Cross-Dataset Bloom Question Classification: Supervised Models and Prompted LLMs
  • Edumirror Educational Social Dynamics — EduMirror: Modeling Educational Social Dynamics with Value-driven Multi-agent Simulation
  • AI Engineering Education Balancing Act — Using AI in engineering education: a balancing act, driven by clear purpose
  • LLM Sentiment Analysis Education Research — LLM-assisted sentiment analysis for integrated computational and qualitative mixed methods education research: A case study of students' written reflection assignments
  • Shame Guilt AI Regulation Computing Education — Stuck in a Spiral": Shame and Guilt as Social Regulators of AI Use in Computing Education
  • Evaluating Interactivity Automated Assessment AI Generated Explorable Explanations — Evaluating Interactivity: Toward Automated Assessment of AI-Generated Explorable Explanations
  • Citation

    Markos Stamatakis, Omkar Gavali, Joshua Berger, Christian Wartena, Anett Hoppe, Ralph Ewerth (2026). Are LLM-based Chatbots Good Enough to Support Computer Science Students in Multiple-Choice Exercises?. arXiv:2606.15919. arXiv preprint.