On this page

CS Education — computer science education is the most-researched STEM subfield in the knowledge base, benefiting from natural alignment between AI tools and programming tasks. Code generation, debugging assistance, and automated code review are its primary AI applications. Because students learn to build the very tools they use, CS education sits at the center of debates about AI literacy, curriculum redesign, agentic software engineering, and the boundary between genuine learning and over-reliance.

Questions to Consider

  • If AI can now write code that passes real programming exams, what should students still learn to do by hand — and what should the curriculum stop teaching?
  • Research found that higher trust in an AI coding assistant predicted WORSE ability to tell correct from misleading suggestions. How is trust different from appropriate reliance, and how would you teach the latter?
  • In student-AI co-programming, nearly 80% of interactions relied on non-learning strategies like outsourcing answers, and only about 1 in 9 showed deep epistemic engagement. Why does genuine learning rarely happen by default when AI is available?
  • A Learning by Teaching agent that was too competent undermined students' debugging practice. Would you deliberately make an AI tutor fallible — and if so, how?
  • As AI automates implementation, curricula are shifting from writing code to verifying and directing AI-generated artifacts. What new competencies does that demand, and what might get lost in the shift?
  • Students are building the very tools they use. How does being both builder and user of AI change what they should learn about its limits — and its ethics?

Introduction

AI in CS education

  • Code generation and completion: CS1 code review, DURA for CS2, and NL programming mistakes examine how students use AI for code generation and what they learn from it.
  • Conversational agents for novices (scoping review): Barzanji & Loitsch (2025) map 23 studies (2019–June 2024) of conversational agents for novice programmers, documenting a shift from rule-based chatbots to Large Language Models (LLMs)- and RAG (Retrieval-Augmented Generation)-based agents (with retrieval-augmented generation reducing hallucination) and personalized tutoring support (e.g., InfoBot, ProbSol-Bot, Lint Bot, Profe Alex). Notably, only 4 of 23 studies ground design in learning theory, and 17 of 23 prototypes are English-only despite most research originating outside English-speaking countries — flagging weak pedagogical grounding and an inclusivity gap for future CA design in introductory programming.
  • Debugging support: Debugging tools, human-AI debugging collaboration, and dyadic pair-programming modeling leverage AI for error identification and repair.
  • Automated assessment: Linux Bash grading, a large-scale 18-model grading comparison, and LLM intervention review evaluate automated code assessment. The security of that workflow is a separate question: Humble (2026) red-teamed a routine AI grading task and found that instructions hidden inside a submitted file raised a failing essay's grade with no visible warning, in 9 of 9 iterations for one strategy and 17 of 18 for another — evidence that grader robustness belongs on the Assessment Validity checklist alongside accuracy.
  • AI-generated learning media: Generated Animated Traces show that AI-generated visualizations can aid immediate learning but must be personalized — mid-engagement students experienced a performance decrement consistent with the expertise-reversal effect.
  • GenAI analogy critique as an instructional asset: Bernstein & Sibia (2026) ground GenAI analogy reception in CS2: ten students who had already completed CS2 audited GenAI-generated analogies for linked lists and recursion and rejected mappings that failed structural correspondence — an island-route analogy that mapped to a circular rather than a singly linked list, or a badminton rally offered for recursion despite having no guaranteed shrinking input, with one student proposing golf instead. The work argues that analogy critique is itself a check on concept understanding, making flawed AI analogies a usable instructional asset rather than a hazard to filter out, and recommends assigning them as objects to inspect and repair.
  • Misconception modeling: a taxonomy of conditionals/loops misconceptions gives automated systems a precise vocabulary for diagnosing novice errors.
  • Model-generated strategic misconceptions: Miličević et al. (2026) built SocraticTrap-CS around 35 concepts from the ACM/IEEE CS2023 curriculum — algorithms, programming languages, databases, networks and operating systems — and asked seven open-weight models for a fluent, authoritative explanation resting on a subtle error. Six of the seven produced an expert-confirmed strategic misconception for 91% or more of prompted concepts (221 of 241 segments, 91.7%), with no significant differences between CS domains; conceptual errors dominated (66.5% conceptual vs. 33.5% factual, and none purely logical), and both persuasiveness and error type varied by domain. The authors therefore recommend domain-sensitive countermeasures — reasoning-focused checks in programming-heavy courses, cross-referencing against protocol specifications in networking — and evaluation of Automated Question Generation and AI-authored explanations on pedagogical trustworthiness rather than correctness alone.
  • Authentic Assessment performance: Lepp & Kaimre (2026) show 2026 GenAI systems outscore the average student cohort on authentic introductory OOP assessments and frequently earn full marks on longer programming tasks, yet still struggle with interfaces, abstract classes, inheritance, and image-based questions — recurring error patterns instructors can exploit when designing assessments.
  • Predictive modeling for at-risk support: Zhang, Jeffries & Koprinska (2025) show that intrinsically interpretable decision trees trained on content-interaction log features accurately predict module-level progress in large-scale online programming courses (85–91% accuracy across four K-12 courses) and flag "No submission" dropout outcomes, giving educators a 7–8 day window to intervene with struggling and disengaged Learners before module deadlines — complementing the automated-grading and attrition-prediction work above.

Programming pedagogy: from blocks to embodied, game-based learning

Programming education spans introductory block-based programming to advanced software development, and increasingly grounds abstract code in concrete, observable outcomes.

  • Block-based visual programming: environments like Scratch and Blockly let beginners snap together graphical blocks rather than type text, eliminating syntax errors and making program structure visible — especially valuable for younger learners and for controlling educational robots. In the AI era they are increasingly combined with conversational AI agents (e.g., Micro:bit + MakeCode in teacher training, small-language-model tutors).
  • Embodied block programming: RoboBlockly Studio combines block-based programming with a conversational AI teaching agent and embodied robot execution, creating an iterative authoring–running–observing–revising loop that preserves learner Learner Agency.
  • Natural-language robot control: EduSim-LLM lets beginners control simulated robots through natural-language instructions, lowering the barrier to robot programming without requiring low-level code expertise.
  • Robotics and computational thinking: Valls i Pou links computational thinking to educational robotics in secondary STEAM curricula, and teacher-training research argues robotics and ML activities should be embedded in Professional Development.
  • Game-based and gamified learning: A systematic review compares game-based learning (suited to informal settings) and gamification (suited to formal classrooms) in robotics education, which emphasizes introductory programming and modular kits.
  • Project-based robotics: Bots and Blocks teaches robotics programming through an agile, semester-spanning project, addressing the theory-practice gap in higher-ed.
  • LLM impact on learning outcomes: Jošt et al. (2024) and a meta-analysis of GenAI and programming learning examine whether AI-assisted tools help or undermine programming achievement.

Curriculum transformation in the AI era

The question "what should students still learn by hand?" now reshapes computing programs.

  • From implementation to verification: Reshaping Undergraduate CS Education argues that as GenAI automates implementation-level programming, debugging, and testing, curricula must shift toward understanding and verifying AI-generated artifacts, preserving system design, abstraction, and critical evaluation while de-emphasizing low-level implementation details. This aligns with AI Literacy frameworks that prize evaluation over generation.
  • Agentic software engineering as a discipline: ASE-26 formalizes directing agents rather than writing code — teaching auditability, context engineering, verification, multi-agent workflows, and AgentOps — and positions agentic AI competence as a structured, scaffolded curriculum rather than syntax mastery.
  • New pedagogies and assessment models: Test-Driven AI-Assisted Learning replaces lectures with self-directed AI-assisted study gated by weekly closed-book tests, preserving individual accountability while AI agents scale material production and marking under human oversight.
  • What predicts Vibe Coding success — and what to keep teaching: A preregistered CHI 2026 study (N=100) of pure "no-code" vibe-coding found that both computer-science achievement (r = .39) and written-communication proficiency (r = .29) independently predicted performance, with CS achievement remaining significant even after controlling for domain-general cognitive ability and contributing roughly twice the unique variance of writing skill. Because the environment hid generated code, CS knowledge could only help indirectly (problem decomposition, algorithmic thinking) — making the CS estimate a lower bound for AI-assisted workflows that also permit editing. The authors argue curricula should weigh written communication alongside CS fundamentals, rather than treating vibe coding as syntax mastery made obsolete.

AI literacy, agency, and the risk of over-reliance

Because programming is where AI assistance is most powerful, it is also where the failure modes are most visible.

  • Trust ≠ appropriate reliance: Trust and reliance on AI (Pitts et al.) find that higher trust in an AI assistant predicted worse discrimination between correct and misleading suggestions during Python Problem Solving — moderated by AI Literacy and need for cognition. Calibration, not confidence, is the goal.

  • Epistemic AI literacy: Wu (2026) shows that in student-AI co-programming, 78.8% of interactions relied on non-mastery-oriented aims and unreliable strategies (outsourcing, verification-seeking), with only 11.1% showing high epistemic engagement — genuine learning rarely emerges without deliberate design support.

  • Structural interventions against copy-paste over-reliance: Soft barriers for copying in AI-assisted programming evaluate lightweight design interventions (e.g., mechanisms that discourage blind copy-paste of AI output) and find they can reduce over-reliance without blocking AI assistance — evidence that the over-reliance risk in CS education is amenable to instructional-design fixes, not just learner-education or bans.

  • Teachable agents and productive practice: Learning-by-teaching with ChatGPT improved knowledge gains and code quality but undermined error-correction practice because the agent is too competent — a design lesson: make agents deliberately fallible so debugging is preserved.

  • Assistance governance: a scoping review of 90 systems introduces the PEA framework (Policy, Enforcement, Authority) for bounding and controlling LLM assistance — a comparative vocabulary for designing Scaffolding that limits over-reliance.

  • Behavioral context for adaptive AI tutoring: Barron et al. (2026) present TutorTrace, a dataset and pipeline that makes learners' behavioral context computable in real time from IDE telemetry in AI-assisted Python courses (N=480). It derives a taxonomy of activity before, between, and across AI queries, and can classify whether an upcoming query reflects guided or dependent Help-Seeking (AUROC=.717) and predict imminent queries (AUROC=.726); behavior-aware prompts reduced no-independent-work query intervals from 50.0% to 20.7% in a preliminary evaluation. This shows how behavioral telemetry can make AI programming tutors adaptive to learners' actual effort, not just their explicit requests.

  • The duality of building what you use: CS students' unique position creates both meta-cognitive awareness of AI limitations and real risk of Over-Reliance on AI-generated code. Code review interviews and critical engagement studies address this tension directly.

  • Agentic coding and comprehension in team PBL: Tanaka et al. (2026) introduced Spec-Driven Development with AI agents into an undergraduate software-engineering project course and found that implementation throughput (added LOC) rose across 2022-2025 while heavy AI use coincided with code-comprehension dips that only recovered after one-on-one instructor checks - direct evidence that throughput gains do not guarantee understanding, and that over-reliance in AI-assisted coding is amenable to instructional monitoring.

Equity, culture, and who gets into computing

  • Broadening participation: SuaCode documents motivations for smartphone-based coding among African students (fewer than 1% of secondary-school leavers have fundamental coding skills), informing accessible AI-supported MOOCs for low-resource contexts.
  • Neurodivergence and collaboration: Neurodivergent computing students report discomfort with ambiguous collaboration structures; structured assignments, smaller consistent teams, and explicit roles improve Accessibility — design lessons for the AI tools entering computing classrooms.
  • Culture shapes perceived ethics: Canadian vs. South Korean computing students judged identical AI-assisted coding practices differently despite functionally identical policies — policy harmonization does not produce perception harmonization, an Academic Integrity and Equity concern.
  • Collaboration transparency: Graf et al. find that partners' misaligned beliefs about each other's AI use predict lower project scores, especially for lower-performing students — transparency mechanisms (disclosures, shared logs) may be needed in collaborative programming.

Ethics education and the workforce

Connections

CS education connects to Computational Thinking, STEM Education, Automated Grading, Prompt Engineering, AI Literacy, Agentic AI, Curriculum Design, Human AI Collaboration, Higher Education, K-12, and Workplace Learning. Its closest applied neighbour is Information Technology Education: CS education takes the program and the algorithm as its object, whereas IT education takes the deployed organizational system and the practitioner's judgment about it — which is why the two fields debate different AI harms, whether code generation erodes programming skill against whether AI troubleshooting erodes diagnostic skill, and ask different questions of their graduates, whether they can build a system against whether they can govern the systems they administer. It is the domain where AIED tools are both used and built, making it a testbed for Intelligent Tutoring, Robots in Education, Collaborative Learning, Game-Based Learning, and the risks of Over-Reliance.

A 72-study synthesis and the VIE framework. Kumar, Wongsirichot and Nanthaamornphong (2026) reviewed the empirical literature on generative AI in computing and programming education (January 2022 – April 2026, 72 studies, 33 venues) and foreground exactly the structural feature that makes the discipline distinctive: the AI generates the assessable artifact itself, so using the tool, learning the skill and being assessed collapse into one keystroke. Their synthesis of 14 themes finds the field's most replicated effect — short-term efficiency and completion gains (36 studies) — is also its most misleading: those gains do not transfer to unaided performance (21 studies), and prior knowledge moderates whether assistance becomes durable skill or a crutch. Detection research is thin (3 studies) while course redesign is comparatively well evidenced (25 studies), and the review consolidates the corpus into three interdependent design requirements — Verification, Implementation and Equity — where critical engagement with AI output must be a graded, observable component of student work rather than an aspiration left to student discretion (Scaffolding, Assessment Validity).

Implications for computing instructors

  • Design assessments AI cannot coast through. Exploit GenAI's recurring failure patterns (interfaces, abstract classes, inheritance, image-based tasks) instead of banning tools outright — GenAI systems still struggle there.

  • Calibrate trust, don't just build it. Trust-reliance research shows higher trust predicted worse discrimination of misleading AI suggestions; teach verification and critical evaluation, moderated by AI literacy and need for cognition.

  • Keep debugging and productive struggle alive. Choose tools or deliberately fallible agents (learning-by-teaching) that preserve error-correction practice, and personalize AI-generated media to avoid expertise-reversal effects (expertise-reversal).

  • Govern AI assistance explicitly. Define policy, enforcement, and authority for LLM support (PEA) rather than leaving boundaries implicit.

  • Shift curricula toward verification and agent direction. As GenAI automates implementation, teach understanding/verifying AI artifacts (reshape curricula) and structured agentic-software-engineering skills (ASE-26).

  • Structure collaboration for all learners. Smaller consistent teams, explicit roles, and AI-use transparency support neurodivergent students and fair collaboration, especially where misaligned AI-use beliefs lower project scores.

  • LLM-adaptive explanations of programming errors (2026): A crowdsourced study (N=103) found LLM-rewritten error messages improve readability, but objective debugging performance depends on matching explanation style (pragmatic vs contingent) to programmer skill — a scaffolding insight for AI-assisted programming education (Beyond the Traceback: Using LLMs for Adaptive Explanations of Programming Errors).

Connected Concepts

Connected Articles

Embed this page

Copy the code below to embed a chromeless version of this page in a learning management system or other website. The embedded view hides the site header, navigation, and footer.