Research Article
Understanding Student Perceptions, Mistakes, and Debugging Approaches when Solving Natural Language Programming Tasks
Synthesis: Learning to communicate with code-generating AI is an emerging skill for novice programmers. 'Prompt Problems' — having students solve computational tasks by writing natural-language prompts for code-generating models — is a recent pedagogical approach, yet little was known about the specific prompt-level mistakes novices make, the computational details they fail to communicate, and how they recover when generated code is wrong. Padurean et al. (2026) studied attempts by more than 900 students to solve dialogue-based Prompt Problems in a CS1 course, analyzing the Misconceptions about AI and repair strategies that surface when learners must specify intent in English rather than code. The study extends the Teaching Prompt-Based Programming with LLMs: A 45-Minute Lesson with Guided Practice for End-User Programmers lineage and the broader Reshaping Undergraduate Computer Science Education in the Generative AI Era movement, situating prompt-writing as a core AI Literacy competency within CS Education. It also connects to Programming Intelligent Tutoring Systems (where natural-language specification has long been a goal) and highlights the need for Scaffolding that helps novices articulate computational detail. Findings on Student Experience and recovery behavior inform Higher Education course design as Large Language Models (LLMs) pair-programming becomes routine.
What this means for practice
- Instructors. Teach constraint specification as a syllabus item, not an incidental skill: 675 students in the first batch and 821 in the second generated at least one initially incorrect prompt that required clarification.
- Instructors. Teach a repair vocabulary for wrong generated code — reframe the task, clarify a constraint, update a signature, add an edge case — so students do not respond to failure by simply re-asking the same prompt.
- Instructors. Build in reflection on the debugging process, since perceived difficulty and frustration tracked the model's difficulty in interpreting intent rather than the task itself.
- Researchers. Analyze iterative prompt refinement against students' self-reported strategies; the study stopped short of linking the two, which is the open question for advanced courses where test cases are not provided.
- Administrators. Fund Scaffolding around prompt construction for CS1 at scale: with more than 900 students attempting these tasks, prompt-level mistakes are the common case, not the exception.
Limitations
- The study is confined to an introductory C programming course and its relatively simple problems, so findings may not generalize to other languages or to realistic, advanced programming scenarios.
- Student demographic data could not be collected under the terms of the ethics approval, so whether the approach is equitable across student populations remains untested.
- Cognitive load was not measured, leaving the alignment with cognitive load theory theoretical rather than empirical.
- The analysis describes prompt mistakes and reported debugging strategies without an in-depth study of how students iteratively refine prompts or whether that refinement improves outcomes.
Citation
Victor-Alexandru Padurean, Kaitlin Riegel, Gweneth Barbre, Musa Blake, Paul Denny, Adish Singla (2026). Understanding Student Perceptions, Mistakes, and Debugging Approaches when Solving Natural Language Programming Tasks. [cs.CY].