Research Article
Exploring the Effectiveness of Using LLMs for Automated Assessment of Student Self Explanations in Programming Education
Synthesis: This paper presents a rigorous empirical comparison between LLM-based and semantic similarity methods for automated assessment of student self-explanations in programming education. The task is framed as binary classification — determining whether a student's explanation of a worked-example step is correct or incorrect.
Worked examples — step-by-step problem solutions — are a well-established Scaffolding technique, and their effectiveness increases when students are prompted to self-explain each step. However, manually assessing these self-explanations doesn't scale. The prevailing approach has been to compare student responses to reference explanations using semantic similarity metrics, but recent advances in large language models raise the question of whether LLM-based scoring now outperforms these traditional methods.
The authors address a critical gap: high-quality, domain-specific datasets with balanced class distributions for automated scoring tasks. Their contribution is both methodological (a rigorous comparison framework) and empirical (which approach works better, and under what conditions).
- Binary classification framing: Self-explanations scored as correct or incorrect, a practical framing for real-world deployment in intelligent tutoring systems
- Dataset contribution: Domain-specific labeled data for programming self-explanations with balanced classes
- Method comparison: LLM-based scoring versus semantic similarity methods, with systematic evaluation
- Practical implications: Guidance for building automated feedback systems in programming education
Connection to Knowledge Base
This work extends the Automated Grading landscape by addressing a specific gap: assessment of open-ended self-explanations rather than final answers or code submissions. It complements research on Confidence Estimation in Automatic Short Answer Grading with LLMs and The Effects of Structured LLM-Generated Feedback on Programming Assignment Performance by focusing on the formative, metacognitive dimension of student learning rather than summative evaluation.
What this means for practice
- Software developers. Score open-ended self-explanations with an LLM-as-judge rather than semantic similarity: LLM Behavior reached F1 = 0.98 and accuracy = 0.96, against F1 = 0.69 and accuracy = 0.6 for Deep Tutor and F1 = 0.7 and accuracy = 0.61 for RoBERTa.
- Software developers. Define correctness in the prompt and match it to how students actually explain code: the construction-based variant underperformed (F1 = 0.89, accuracy = 0.8) because the student explanations described "syntax and code behavior" and omitted the context of the line within the rest of the program.
- Designers. Optimize automated scoring for false negatives when students see the result, since the best LLM Behavior fold on the test sets produced TP=357, FP=5, FN=1, TN=2, and a correct explanation scored incorrect is what frustrates students who know they are right.
- Instructors. Check class balance before trusting a self-explanation benchmark: the SelfCode2.0 subset used here held 1794 correct explanations against only 60 incorrect ones, so the balanced test set had to be completed with synthetically generated negative examples.
Limitations
- Only 60 of the 1794 labeled student explanations were incorrect, so the balanced dataset was completed with synthetic negatives generated by GPT-OSS 20b rather than real student errors, and those synthetic items stand in for the negative class in every reported metric.
- The evaluation is a zero-shot benchmark with GPT-3.5-Turbo-16k on the SelfCode2.0 dataset — 60 students and 2 experts over four code examples — in one domain, so results may not transfer to other models or disciplines.
- Scoring is binary correct or incorrect at the level of a single code line, which cannot represent partially correct explanations or distinguish minor from severe errors.
- No student learning outcome was measured; the study establishes scoring agreement, not whether automated feedback improves comprehension.
Citation
Lekshmi-Narayanan, A.-B., Hassany, M., & Brusilovsky, P. (2026). Exploring the Effectiveness of Using LLMs for Automated Assessment of Student Self Explanations in Programming Education.