On this page

Synthesis: Bernstein, Denny, Leinonen et al. (2026) investigate whether providing students with multiple, diverse LLM-generated explanations of code (rather than a single 'best' explanation) improves comprehension in introductory programming. Their findings show that exposure to diverse explanations significantly outperforms single-explanation conditions on measures of conceptual understanding and code comprehension. This challenges the common design assumption that AI-generated educational content should converge on a single 'correct' explanation, instead suggesting that Large Language Models (LLMs)-generated Feedback Loop diversity supports deeper learning by exposing students to multiple perspectives. The study connects to Scaffolding theory, where multiple representations support the gradual transfer of responsibility from tool to learner. It also informs Active Learning Pedagogies and Teaching Strategies by providing a concrete implementation strategy for AI-assisted instruction. The work has implications for how Student Experience of programming education can be enhanced through deliberately varied AI-generated content, relevant to STEM Education course design.

What this means for practice

  • Instructors. Give students several explanations that each emphasize a different dimension — function, concept, goal — instead of repeating one general-purpose explanation. Open-ended response accuracy was consistently about 7.7% higher in the diverse condition.
  • Instructors. Do not assume that extra explanations cost working memory: perceived cognitive load did not differ between the diverse and generic conditions.
  • Designers. Generate each explanation with its own dimension-targeted prompt rather than asking the model for a single best explanation, and offer students the angle they find most helpful; this supports differentiated instruction at scale and matches how students already use generative AI tools.
  • Instructors. Treat diversity as a promising pattern, not a proven intervention. The comprehension gains were not statistically significant, and high overall scores point to a possible ceiling effect.
  • Researchers. Size future studies to detect small effects and test retention, since students saw only two short exercises and were tested immediately after.

Limitations

  • All 971 first-year computing students came from a single course taught by the same instructor, and prior knowledge of recursion was not measured, so condition differences may reflect unmeasured preparation.
  • Students received only two short programming exercises (sumArray plus either randomizeString or countChar) and were tested immediately, which limits the magnitude of any observable effect and may capture short-term rather than durable learning.
  • High overall scores point to a plausible ceiling effect in which small improvements are hard to detect.
  • The study cannot verify whether or how thoroughly students engaged with the explanations, the three explanation dimensions were never compared individually, and students had no chance to ask follow-up questions — closer to an instructor distributing curated written explanations than to a live LLM deployment.

Citation

Seth Bernstein, Paul Denny, Juho Leinonen, Kush Patel, Rayhona Nasimova, Matt Littlefield, Stephen MacNeil (2026). Exploring the Value of Diverse LLM Explanations in Introductory Programming. cs.HC (SIGCSE Virtual 2026).

Embed this page

Copy the code below to embed a chromeless version of this page in a learning management system or other website. The embedded view hides the site header, navigation, and footer.