On this page

Synthesis: This paper investigates whether large language models can serve as scalable proxies for students by simulating realistic logical errors in code submissions. Using the CodeWorkout dataset of 74,000+ unique student Java submissions across 37 problems, the authors evaluate five LLMs under three prompting strategies: Input-Output (IO), Chain-of-Thought (CoT), and iterative Self-Refine.

  • Diversity vs. Alignment trade-off: LLMs generate diverse error patterns, but alignment with authentic student errors varies significantly by model. Claude Sonnet 4 achieves the most balanced performance across both dimensions.
  • Functional indistinguishability: A blinded expert annotation study (N=401) found that synthetic errors are functionally indistinguishable from authentic student errors.
  • Task difficulty effects: Higher-struggling-level problems elicit more diverse but less student-like errors — LLMs struggle more to simulate realistic mistakes on harder tasks.
  • Practical implications: Synthetic errors could be integrated into intelligent tutoring systems, teachable agents, and large-scale learning analytics pipelines without waiting for authentic classroom data accumulation.

Methodology

The study used the CodeWorkout dataset with 74,000+ unique student Java submissions. Five LLMs were tested under three prompting strategies. Performance was assessed on two dimensions: diversity (range of distinct error patterns) and alignment (correspondence with authentic student mistakes). A blinded expert annotation study with 401 samples confirmed the indistinguishability of synthetic and authentic errors.

This work extends research on LLM-based student simulation and student misconception identification. It connects to programming intelligent tutoring systems and student modeling by offering a scalable method for generating training and evaluation data. The findings also inform AI-generated traces from novice programmers and research on code review with generative AI in CS1.

What this means for practice

  • Learners. Rehearse debugging on simulated mistakes rather than waiting for real ones: errors generated by large language models were functionally indistinguishable from authentic student errors, so diagnosing an AI proxy's code is a workable way to train error-spotting (Simulating Students, Learning by Teaching).
  • Learners. Drill the fault types that dominate novice submissions — synthetic bugs concentrated on Condition Logic (46.4%) and Boundary (21.9%) errors — and expect that harder problems yield more diverse but less student-like mistakes.
  • Learners. Do not treat plausibility as a source cue: annotators misclassified 164 of 196 (83.7%) LLM-generated submissions as human-written, and the synthetic errors scored higher on plausibility (M = 4.27, SD = 1.07) than authentic student errors (M = 3.78, SD = 1.11).
  • Learners. Use synthetic errors to isolate one misconception at a time, and keep authentic code in your practice mix: the authors describe the model as an "idealized learner" that is a less faithful proxy for the full messiness of authentic student work.

Limitations

  • The experiments are grounded in CodeWorkout and a curated set of 37 introductory Java problems, so the observed diversity–alignment trade-offs may not generalize to other institutions, curricula, assessment formats, or programming languages.
  • Generation constraints required compilable code with exactly one non-trivial logical error, which reduces the syntactic noise and multi-fault behavior common in authentic novice submissions and produces an "idealized learner" distribution.
  • The blinded annotation set was single-coded: the 401 items (205 authentic, 196 synthetic) were split between two annotators with no formal inter-rater reliability statistic, with calibration resting on a pilot of 43 submissions.
  • Struggling level was operationalized as the total number of submissions per problem, which the authors note may conflate difficulty with assignment placement, popularity, or course policies, and the AST edit distance they use as the diversity proxy is purely structural — functionally equivalent code can have distinct AST structures.

Citation

Keramati, A., Cao, J., Mohammadi, I., Warschauer, M., & Shi, Y. (2026). Simulating Students' Java Programming Errors with Large Language Models.

Embed this page

Copy the code below to embed a chromeless version of this page in a learning management system or other website. The embedded view hides the site header, navigation, and footer.