Research Article
Towards SocratiCode: Designing a Generative AI-Based Programming Tutor for K-12 Students through a 4-Week Participatory Design Study
Synthesis: Socratic questioning, reflection prompts, misconception checks, and mandatory pauses produce better K-12 engagement than directive answer-giving AI tutors.
Synthesis
SocratiCode demonstrates a participatory design evolution from a directive AI tutor to a Socratic learning companion for K-12 programming instruction. Over four weeks with two K-12 Python learners, the system shifted from flexible tutorial generation toward dialogic support: guided questioning instead of answers, reflection prompts, misconception checks, incremental hints, and mandatory pauses requiring learner input. This Socratic shift improved explanation clarity and Problem Solving engagement. The findings directly reinforce a Socratic discovery-based approach over direct answer generation, but extend it to the K-12 context where cognitive load concerns are particularly acute. The emphasis on mandatory pauses and reflection aligns with Metacognition and Self-Regulated Learning Scaffolding strategies. The authors argue that AI tutoring is most effective as a companion within a human-guided framework, not an answer engine — a principle that resonates with the Human-in-the-Loop architecture and the findings from The Effects of Structured LLM-Generated Feedback on Programming Assignment Performance that less guided feedback may be more effective.
Key Findings
- Adaptive Generative AI that shifts from directive tutorial generation toward Socratic questioning—guided questioning, reflection prompts, misconception checks, incremental hints, and mandatory pauses—improves explanation clarity and sustained engagement for K-12 beginners.
- Across four weekly participatory iterations the prompt evolved from an explanatory tutor into a stabilized, dialogic flow defined by hard stopping points, learner-response requirements, and no unsolicited next steps.
- Learners consistently preferred incremental hints and multiple attempts over immediate full solutions, explicitly requesting more guided hinting rather than answers.
- Human oversight remained essential—particularly for correcting Misconceptions about AI, challenging weak reasoning (including the system acting as a “yes-man”), and supporting advanced topics like recursion and functions.
Background and Motivation
Generative AI large language models such as ChatGPT and Claude can produce explanations, worked examples, and step-by-step tutorials on demand, opening new paths for scalable programming instruction. Yet many existing systems remain answer-oriented: they rarely determine appropriate stopping points, often ignore the learner’s prior expertise, may introduce out-of-scope concepts, and can generate content misaligned with learner goals, creating confusion rather than understanding. In programming education these limitations are especially harmful, because novices need opportunities to articulate their reasoning, reflect on misunderstandings, and work through problems before receiving complete solutions.
These challenges intensify in K-12 contexts, where learners have little prior exposure and struggle comparatively more with abstract reasoning, pacing, and engagement. Prior work by novice-programmer researchers (e.g., overreliance on ChatGPT, hallucinations, weak rationales, and difficulty interpreting AI-generated code) shows that novices become dependent on external solutions unless the system supports debugging and active reasoning. The paper argues that K-12 generative AI should support scaffolding, debugging, and active reasoning rather than simply produce code or tutorial-style explanations.
SocratiCode Design
SocratiCode is an adaptive, prompt-based tutorial system built on established Adaptive Learning research, generative AI in education, and Socratic tutoring principles. The model was instructed to collect background information from learners before generation, assume a beginner user, and default to Python unless requested otherwise. The prompt was structured into multiple high-level components — System Role; Learner Level and Background Selection; Tutorial Structure and Flow Control; Reinforcement, Adaptivity, and Closure; and Constraints and Content Boundaries — so that generated lessons embedded introductions, examples, practice, summaries, and follow-up tasks that created natural stopping points.
To keep instruction developmentally appropriate in K-12 settings, the system adjusted pacing, analogies, and explanations to individual learners. Later iterations layered in misconception clarification, metaphor use, reflective pauses, and hints provided before complete solutions, moving the system from content delivery toward guided inquiry. The final template defined a lesson flow of hook or analogy → concept explanation → code walkthrough → short exercise → optional misconception note → reflection → transition, with interaction rules requiring a pause after exercises and learner input before proceeding.
Four-Week Participatory Design Process
The framework was developed through a participatory design study with two Grade-11 high school interns (one male, one female, both 17–18, with no prior programming experience) during a summer 2025 internship at a university. The first prompt version was deployed on the GPT platform with GPT-5 as the default model. A four-week curriculum covered fundamental topics aligned with the ACM/IEEE Computer Science Curricula (CS2013) introductory programming guidelines: variables and conditionals in W1, progressing to loops, arrays, and functions by W4. A master’s-student teaching assistant introduced each topic and supplied 3–4 practice problems, and a computer science faculty member ran one-to-one and group meetings.
Feedback collection followed an agile-style mixed-methods protocol: daily stand-ups combined closed-ended questions (task completion, prior exposure, expert consultation) with open-ended prompts about what was understood and confused, while weekly sync meetings added 5-point Likert surveys and open-ended reflections. Two authors independently conducted open coding and thematic analysis using an inductive approach, yielding four key categories: Engagement and Appeal, Human–AI Collaboration, Explanations and Clarity, and Instructional Design and Structure. Prompt revisions were introduced at the start of W2; W3 and W4 designs were informed by prior-week feedback; by W4 no further modifications were needed, indicating the prompt had stabilized.
Findings
Four themes emerged from daily and weekly learner feedback. First, engagement was strengthened when the system adapted explanations to learners’ prior knowledge, pace, and responses, with participants praising relevant examples and anecdotes that built on what they already knew. Second, definitions, checkpoints, and staged pacing improved conceptual clarity — clarity dropped when new topics advanced without defining key terms, and breaking concepts into smaller steps linked to familiar contexts aided comprehension. Third, learners preferred incremental hints and multiple attempts before full answers, though immediate or inconsistent explanations sometimes caused confusion. Fourth, human oversight remained necessary, especially for more advanced or context-dependent material; participants regarded expert assistance as still essential, and noted that SocratiCode sometimes acted as a "yes-man," prioritizing its own answers over learner reasoning.
What this means for practice
- Learners. Ask for the hint, not the answer, and expect to attempt a problem more than once: both participants consistently preferred incremental hints and multiple attempts over immediate full solutions and explicitly requested more guided hinting.
- Learners. Treat the pause as part of the work. The stabilized design required learner input before proceeding and used hard stopping points after each exercise rather than a continuous stream of new material.
- Learners. Do not let the tutor's confidence settle a question for you: participants reported that SocratiCode sometimes acted as a "yes-man," prioritizing its own answers over learner reasoning, and one participant resolved confusions about Python 0-indexing and mistaking
=for==only with expert intervention. - Instructors. Budget expert time alongside the tool rather than replacing it. Participants put the need at "two to three times human assistance per week," and human support was what corrected Misconceptions about AI, challenged weak reasoning, and carried advanced topics such as recursion and functions.
- Designers. Encoding Socratic behavior in the prompt structure is what produced the change: structuring it into System Role; Learner Level and Background Selection; Tutorial Structure and Flow Control; Reinforcement, Adaptivity, and Closure; and Constraints and Content Boundaries, with a pause required after exercises, moved the system from content delivery toward guided questioning.
Limitations
- The study rests on two participants — grade 11 interns, one male and one female, between 17 and 18 years old with no prior programming experience — and the authors state this limits the generalizability of the findings.
- It evaluated a single customized GPT-based system with GPT-5 as the default model, so other generative AI models may behave differently under similar conditions.
- The content was Python only, following ACM/IEEE CS2013 introductory guidelines from variables and conditionals in W1 to loops, arrays, and functions by W4, which the authors note may limit transferability to other programming languages.
- Evidence is self-report: daily stand-ups and weekly meetings with a 5-point Likert survey and open-ended reflection, coded by two of the authors, so the claims concern perceived clarity and engagement rather than measured programming gains, and the underlying model may reflect biases in training data, representation, and instructional style.
Citation
Lucas, C., Tsai, C.-H., Bihani, A., & Sarker, J. (2026). Towards SocratiCode: Designing a Generative AI-Based Programming Tutor for K-12 Students through a 4-Week Participatory Design Study.