Concept
CS Education
CS Education — computer science education is the most-researched STEM subfield in the knowledge base, benefiting from natural alignment between AI tools and programming tasks. Code generation, debugging assistance, and automated code review are its primary AI applications. Because students learn to build the very tools they use, CS education sits at the center of debates about AI literacy, curriculum redesign, agentic software engineering, and the boundary between genuine learning and over-reliance.
Questions to Consider
- If AI can now write code that passes real programming exams, what should students still learn to do by hand — and what should the curriculum stop teaching?
- Research found that higher trust in an AI coding assistant predicted WORSE ability to tell correct from misleading suggestions. How is trust different from appropriate reliance, and how would you teach the latter?
- In student-AI co-programming, nearly 80% of interactions relied on non-learning strategies like outsourcing answers, and only about 1 in 9 showed deep epistemic engagement. Why does genuine learning rarely happen by default when AI is available?
- A Learning by Teaching agent that was too competent undermined students' debugging practice. Would you deliberately make an AI tutor fallible — and if so, how?
- As AI automates implementation, curricula are shifting from writing code to verifying and directing AI-generated artifacts. What new competencies does that demand, and what might get lost in the shift?
- Students are building the very tools they use. How does being both builder and user of AI change what they should learn about its limits — and its ethics?
Introduction
AI in CS education
- Code generation and completion: CS1 code review, DURA for CS2, and NL programming mistakes examine how students use AI for code generation and what they learn from it.
- Conversational agents for novices (scoping review): Barzanji & Loitsch (2025) map 23 studies (2019–June 2024) of conversational agents for novice programmers, documenting a shift from rule-based chatbots to Large Language Models (LLMs)- and RAG (Retrieval-Augmented Generation)-based agents (with retrieval-augmented generation reducing hallucination) and personalized tutoring support (e.g., InfoBot, ProbSol-Bot, Lint Bot, Profe Alex). Notably, only 4 of 23 studies ground design in learning theory, and 17 of 23 prototypes are English-only despite most research originating outside English-speaking countries — flagging weak pedagogical grounding and an inclusivity gap for future CA design in introductory programming.
- Debugging support: Debugging tools, human-AI debugging collaboration, and dyadic pair-programming modeling leverage AI for error identification and repair.
- Automated assessment: Linux Bash grading, a large-scale 18-model grading comparison, and LLM intervention review evaluate automated code assessment. The security of that workflow is a separate question: Humble (2026) red-teamed a routine AI grading task and found that instructions hidden inside a submitted file raised a failing essay's grade with no visible warning, in 9 of 9 iterations for one strategy and 17 of 18 for another — evidence that grader robustness belongs on the Assessment Validity checklist alongside accuracy.
- AI-generated learning media: Generated Animated Traces show that AI-generated visualizations can aid immediate learning but must be personalized — mid-engagement students experienced a performance decrement consistent with the expertise-reversal effect.
- GenAI analogy critique as an instructional asset: Bernstein & Sibia (2026) ground GenAI analogy reception in CS2: ten students who had already completed CS2 audited GenAI-generated analogies for linked lists and recursion and rejected mappings that failed structural correspondence — an island-route analogy that mapped to a circular rather than a singly linked list, or a badminton rally offered for recursion despite having no guaranteed shrinking input, with one student proposing golf instead. The work argues that analogy critique is itself a check on concept understanding, making flawed AI analogies a usable instructional asset rather than a hazard to filter out, and recommends assigning them as objects to inspect and repair.
- Misconception modeling: a taxonomy of conditionals/loops misconceptions gives automated systems a precise vocabulary for diagnosing novice errors.
- Model-generated strategic misconceptions: Miličević et al. (2026) built SocraticTrap-CS around 35 concepts from the ACM/IEEE CS2023 curriculum — algorithms, programming languages, databases, networks and operating systems — and asked seven open-weight models for a fluent, authoritative explanation resting on a subtle error. Six of the seven produced an expert-confirmed strategic misconception for 91% or more of prompted concepts (221 of 241 segments, 91.7%), with no significant differences between CS domains; conceptual errors dominated (66.5% conceptual vs. 33.5% factual, and none purely logical), and both persuasiveness and error type varied by domain. The authors therefore recommend domain-sensitive countermeasures — reasoning-focused checks in programming-heavy courses, cross-referencing against protocol specifications in networking — and evaluation of Automated Question Generation and AI-authored explanations on pedagogical trustworthiness rather than correctness alone.
- Authentic Assessment performance: Lepp & Kaimre (2026) show 2026 GenAI systems outscore the average student cohort on authentic introductory OOP assessments and frequently earn full marks on longer programming tasks, yet still struggle with interfaces, abstract classes, inheritance, and image-based questions — recurring error patterns instructors can exploit when designing assessments.
- Predictive modeling for at-risk support: Zhang, Jeffries & Koprinska (2025) show that intrinsically interpretable decision trees trained on content-interaction log features accurately predict module-level progress in large-scale online programming courses (85–91% accuracy across four K-12 courses) and flag "No submission" dropout outcomes, giving educators a 7–8 day window to intervene with struggling and disengaged Learners before module deadlines — complementing the automated-grading and attrition-prediction work above.
Programming pedagogy: from blocks to embodied, game-based learning
Programming education spans introductory block-based programming to advanced software development, and increasingly grounds abstract code in concrete, observable outcomes.
- Block-based visual programming: environments like Scratch and Blockly let beginners snap together graphical blocks rather than type text, eliminating syntax errors and making program structure visible — especially valuable for younger learners and for controlling educational robots. In the AI era they are increasingly combined with conversational AI agents (e.g., Micro:bit + MakeCode in teacher training, small-language-model tutors).
- Embodied block programming: RoboBlockly Studio combines block-based programming with a conversational AI teaching agent and embodied robot execution, creating an iterative authoring–running–observing–revising loop that preserves learner Learner Agency.
- Natural-language robot control: EduSim-LLM lets beginners control simulated robots through natural-language instructions, lowering the barrier to robot programming without requiring low-level code expertise.
- Robotics and computational thinking: Valls i Pou links computational thinking to educational robotics in secondary STEAM curricula, and teacher-training research argues robotics and ML activities should be embedded in Professional Development.
- Game-based and gamified learning: A systematic review compares game-based learning (suited to informal settings) and gamification (suited to formal classrooms) in robotics education, which emphasizes introductory programming and modular kits.
- Project-based robotics: Bots and Blocks teaches robotics programming through an agile, semester-spanning project, addressing the theory-practice gap in higher-ed.
- LLM impact on learning outcomes: Jošt et al. (2024) and a meta-analysis of GenAI and programming learning examine whether AI-assisted tools help or undermine programming achievement.
Curriculum transformation in the AI era
The question "what should students still learn by hand?" now reshapes computing programs.
- From implementation to verification: Reshaping Undergraduate CS Education argues that as GenAI automates implementation-level programming, debugging, and testing, curricula must shift toward understanding and verifying AI-generated artifacts, preserving system design, abstraction, and critical evaluation while de-emphasizing low-level implementation details. This aligns with AI Literacy frameworks that prize evaluation over generation.
- Agentic software engineering as a discipline: ASE-26 formalizes directing agents rather than writing code — teaching auditability, context engineering, verification, multi-agent workflows, and AgentOps — and positions agentic AI competence as a structured, scaffolded curriculum rather than syntax mastery.
- New pedagogies and assessment models: Test-Driven AI-Assisted Learning replaces lectures with self-directed AI-assisted study gated by weekly closed-book tests, preserving individual accountability while AI agents scale material production and marking under human oversight.
- What predicts Vibe Coding success — and what to keep teaching: A preregistered CHI 2026 study (N=100) of pure "no-code" vibe-coding found that both computer-science achievement (r = .39) and written-communication proficiency (r = .29) independently predicted performance, with CS achievement remaining significant even after controlling for domain-general cognitive ability and contributing roughly twice the unique variance of writing skill. Because the environment hid generated code, CS knowledge could only help indirectly (problem decomposition, algorithmic thinking) — making the CS estimate a lower bound for AI-assisted workflows that also permit editing. The authors argue curricula should weigh written communication alongside CS fundamentals, rather than treating vibe coding as syntax mastery made obsolete.
AI literacy, agency, and the risk of over-reliance
Because programming is where AI assistance is most powerful, it is also where the failure modes are most visible.
-
Trust ≠ appropriate reliance: Trust and reliance on AI (Pitts et al.) find that higher trust in an AI assistant predicted worse discrimination between correct and misleading suggestions during Python Problem Solving — moderated by AI Literacy and need for cognition. Calibration, not confidence, is the goal.
-
Epistemic AI literacy: Wu (2026) shows that in student-AI co-programming, 78.8% of interactions relied on non-mastery-oriented aims and unreliable strategies (outsourcing, verification-seeking), with only 11.1% showing high epistemic engagement — genuine learning rarely emerges without deliberate design support.
-
Structural interventions against copy-paste over-reliance: Soft barriers for copying in AI-assisted programming evaluate lightweight design interventions (e.g., mechanisms that discourage blind copy-paste of AI output) and find they can reduce over-reliance without blocking AI assistance — evidence that the over-reliance risk in CS education is amenable to instructional-design fixes, not just learner-education or bans.
-
Teachable agents and productive practice: Learning-by-teaching with ChatGPT improved knowledge gains and code quality but undermined error-correction practice because the agent is too competent — a design lesson: make agents deliberately fallible so debugging is preserved.
-
Assistance governance: a scoping review of 90 systems introduces the PEA framework (Policy, Enforcement, Authority) for bounding and controlling LLM assistance — a comparative vocabulary for designing Scaffolding that limits over-reliance.
-
Behavioral context for adaptive AI tutoring: Barron et al. (2026) present TutorTrace, a dataset and pipeline that makes learners' behavioral context computable in real time from IDE telemetry in AI-assisted Python courses (N=480). It derives a taxonomy of activity before, between, and across AI queries, and can classify whether an upcoming query reflects guided or dependent Help-Seeking (AUROC=.717) and predict imminent queries (AUROC=.726); behavior-aware prompts reduced no-independent-work query intervals from 50.0% to 20.7% in a preliminary evaluation. This shows how behavioral telemetry can make AI programming tutors adaptive to learners' actual effort, not just their explicit requests.
-
The duality of building what you use: CS students' unique position creates both meta-cognitive awareness of AI limitations and real risk of Over-Reliance on AI-generated code. Code review interviews and critical engagement studies address this tension directly.
-
Agentic coding and comprehension in team PBL: Tanaka et al. (2026) introduced Spec-Driven Development with AI agents into an undergraduate software-engineering project course and found that implementation throughput (added LOC) rose across 2022-2025 while heavy AI use coincided with code-comprehension dips that only recovered after one-on-one instructor checks - direct evidence that throughput gains do not guarantee understanding, and that over-reliance in AI-assisted coding is amenable to instructional monitoring.
Equity, culture, and who gets into computing
- Broadening participation: SuaCode documents motivations for smartphone-based coding among African students (fewer than 1% of secondary-school leavers have fundamental coding skills), informing accessible AI-supported MOOCs for low-resource contexts.
- Neurodivergence and collaboration: Neurodivergent computing students report discomfort with ambiguous collaboration structures; structured assignments, smaller consistent teams, and explicit roles improve Accessibility — design lessons for the AI tools entering computing classrooms.
- Culture shapes perceived ethics: Canadian vs. South Korean computing students judged identical AI-assisted coding practices differently despite functionally identical policies — policy harmonization does not produce perception harmonization, an Academic Integrity and Equity concern.
- Collaboration transparency: Graf et al. find that partners' misaligned beliefs about each other's AI use predict lower project scores, especially for lower-performing students — transparency mechanisms (disclosures, shared logs) may be needed in collaborative programming.
Ethics education and the workforce
- Ethics-to-behavior gap: the "Cost-of-Ethics Crisis" shows CS students, despite contemporary ethics education, prioritize compensation, location, and culture over ethical concerns in job searches — a critical gap in how ethics instruction transfers to behavior.
- Workforce reshaping: a systematic review of U.S. gray literature frames the "Dual Train Problem" — rapid AI change racing institutional adaptation — and urges durable AI competencies, ethics/governance, and skill-based credentials aligned with emerging roles (e.g., Prompt Engineering, AI auditing, AI policy).
Connections
CS education connects to Computational Thinking, STEM Education, Automated Grading, Prompt Engineering, AI Literacy, Agentic AI, Curriculum Design, Human AI Collaboration, Higher Education, K-12, and Workplace Learning. Its closest applied neighbour is Information Technology Education: CS education takes the program and the algorithm as its object, whereas IT education takes the deployed organizational system and the practitioner's judgment about it — which is why the two fields debate different AI harms, whether code generation erodes programming skill against whether AI troubleshooting erodes diagnostic skill, and ask different questions of their graduates, whether they can build a system against whether they can govern the systems they administer. It is the domain where AIED tools are both used and built, making it a testbed for Intelligent Tutoring, Robots in Education, Collaborative Learning, Game-Based Learning, and the risks of Over-Reliance.
A 72-study synthesis and the VIE framework. Kumar, Wongsirichot and Nanthaamornphong (2026) reviewed the empirical literature on generative AI in computing and programming education (January 2022 – April 2026, 72 studies, 33 venues) and foreground exactly the structural feature that makes the discipline distinctive: the AI generates the assessable artifact itself, so using the tool, learning the skill and being assessed collapse into one keystroke. Their synthesis of 14 themes finds the field's most replicated effect — short-term efficiency and completion gains (36 studies) — is also its most misleading: those gains do not transfer to unaided performance (21 studies), and prior knowledge moderates whether assistance becomes durable skill or a crutch. Detection research is thin (3 studies) while course redesign is comparatively well evidenced (25 studies), and the review consolidates the corpus into three interdependent design requirements — Verification, Implementation and Equity — where critical engagement with AI output must be a graded, observable component of student work rather than an aspiration left to student discretion (Scaffolding, Assessment Validity).
Implications for computing instructors
-
Design assessments AI cannot coast through. Exploit GenAI's recurring failure patterns (interfaces, abstract classes, inheritance, image-based tasks) instead of banning tools outright — GenAI systems still struggle there.
-
Calibrate trust, don't just build it. Trust-reliance research shows higher trust predicted worse discrimination of misleading AI suggestions; teach verification and critical evaluation, moderated by AI literacy and need for cognition.
-
Keep debugging and productive struggle alive. Choose tools or deliberately fallible agents (learning-by-teaching) that preserve error-correction practice, and personalize AI-generated media to avoid expertise-reversal effects (expertise-reversal).
-
Govern AI assistance explicitly. Define policy, enforcement, and authority for LLM support (PEA) rather than leaving boundaries implicit.
-
Shift curricula toward verification and agent direction. As GenAI automates implementation, teach understanding/verifying AI artifacts (reshape curricula) and structured agentic-software-engineering skills (ASE-26).
-
Structure collaboration for all learners. Smaller consistent teams, explicit roles, and AI-use transparency support neurodivergent students and fair collaboration, especially where misaligned AI-use beliefs lower project scores.
-
LLM-adaptive explanations of programming errors (2026): A crowdsourced study (N=103) found LLM-rewritten error messages improve readability, but objective debugging performance depends on matching explanation style (pragmatic vs contingent) to programmer skill — a scaffolding insight for AI-assisted programming education (Beyond the Traceback: Using LLMs for Adaptive Explanations of Programming Errors).
Connected Concepts
- Computational Thinking
- Vibe Coding
- STEM Education
- Information Technology Education
- Automated Assessment
- Prompt Engineering
- AI Literacy
- Agentic AI
- Curriculum Design
- Human AI Collaboration
- Higher Education
- K-12
- Robots in Education
- Game-Based Learning
- Generative AI
- Intelligent Tutoring
- Cognitive Offloading
- Professional Development
- Workplace Learning
- AI in Education
- Collaborative Learning
- Prior Knowledge
- Scaffolding
- Assessment Validity
Connected Articles
-
Generative AI in computing education: A systematic review and a framework for responsible integration — Systematic review of 72 studies: efficiency gains that do not transfer, and the VIE framework
-
Computer Science Achievement and Writing Skills Predict Vibe Coding Proficiency — CS achievement and writing skills predict vibe-coding proficiency (CHI 2026)
-
Giving Mechanical Engineers Intelligent Tools: A Project-Based AI Education Curriculum in Thermal Engineering — Project-Based AI Education Curriculum in Thermal Engineering
-
Harnessing Generative Artificial Intelligence in Computer Science Education: Pedagogical Innovation, Ethical Responsibility, and the Future of Assessment — GenAI in CS education
-
Combating Harms of Generative AI in CS1 with Code Review Interviews and a Flipped Classroom — CS1 code review of AI-generated code
-
Demystify, Use, Reflect, Assess (DURA): An Experience Report on LLM Integration in CS2 — DURA: LLM assistants for CS2
-
Reshaping Undergraduate Computer Science Education in the Generative AI Era — reshaping undergraduate CS curricula for GenAI
-
ASE-26: A Curriculum for Agentic Software Engineering as a Discipline — ASE-26 agentic-software-engineering curriculum
-
Test-Driven, AI-Assisted Learning: Replacing Lectures with Weekly Closed-Book Tests — Test-Driven AI-Assisted Learning
-
Revisiting the Performance of Generative Artificial Intelligence on Introductory Object-Oriented Programming — GenAI performance on authentic introductory OOP assessments (Lepp & Kaimre 2026)
-
Trust and Reliance on AI in Education: AI Literacy and Need for Cognition as Moderators — trust vs. appropriate reliance during Python problem-solving
-
Constructing Epistemic AI Literacy: Detecting Epistemic Aims and Processes in Student-AI Co-Programming — epistemic AI literacy in student-AI co-programming
-
Learning-by-Teaching with ChatGPT: The Effect of a Teachable ChatGPT Agent on Programming Education — learning-by-teaching with ChatGPT
-
Exploring the Design Space of LLM-Based Programming Support in CS Education: A Scoping Review through the Lens of Assistance Governance — PEA framework for bounding LLM assistance
-
Exploring Conversational Agents for Novice Programmers: A Scoping Review — Scoping review of conversational agents for novice programmers
-
DebugTracker: Lightweight Process Evidence for Classroom Debugging — DebugTracker classroom debugging
-
A systematic comparison of Large Language Models for automated assignment assessment in programming education: Exploring the importance of architecture and vendor — 18-model automated grading comparison
-
AI-Generated Traces for Novice Programmers: Learning Effects and Learner Differences in a Multi-Institutional Study — AI-generated animated traces
-
How Students (Mis)understand Conditionals and Loops -- A Taxonomy — conditionals/loops misconception taxonomy
-
The Impact of Large Language Models on Programming Education and Student Learning Outcomes — LLM impact on programming learning outcomes (Jošt et al.)
-
A meta-analysis of the effect of generative AI on productivity and learning in programming — meta-analysis of GenAI and programming learning
-
ProPACT: A Proactive AI-Driven Adaptive Collaborative Tutor for Pair Programming — dyadic pair-programming modeling
-
To Tab or Not to Tab: Measuring Critical Engagement in AI Code Completion Tools Using Behavioral Signals and Attention Checks — critical engagement with code completion
-
Why SuaCode?": Understanding African Students' Motivations for Taking a Smartphone-Based Online Coding Course — SuaCode smartphone-based coding in Africa
-
Did Alice Do Wrong? Cross-Cultural Differences in Student Perceptions of Generative AI Use in University Computing Education — cross-cultural perceptions of AI-assisted coding
-
I can't read your mind": A Study of Neurodivergent Computing Students' Experiences with Collaborative Active Learning — neurodivergent computing students
-
Coding, robots, computational concepts, and machine learning using the microbit card and the Maqueen and Nezha kits. A study in initial teacher training — Micro:bit + ML in teacher training
-
Computational Thinking to Enhance Educational Robotics in Secondary School's Curriculum — computational thinking and educational robotics
-
RoboBlockly Studio: Conversational Block Programming With Embodied Robot Feedback for Computational Thinking — RoboBlockly embodied block programming
-
EduSim-LLM: An Educational Platform Integrating Large Language Models and Robotic Simulation for Beginners — EduSim-LLM natural-language robot control
-
Using LLMs to Detect Growth in Computational Thinking in Introductory Physics — LLM support for computational thinking in physics
-
The StudyChat Dataset: Analyzing Student Dialogues With ChatGPT in an Artificial Intelligence Course — The StudyChat dataset of student–LLM dialogues in an AI course
-
Analysis of Types of Inquiries in Student-AI Interaction: A case study of two CS2 tasks — Analysis of Types of Inquiries in Student-AI Interaction
-
ChatGPT Solves All Tested Qiskit Homework Assignments — ChatGPT solves Qiskit homework; autogradable design
-
Beyond the Traceback: Using LLMs for Adaptive Explanations of Programming Errors — LLM adaptive explanations of programming errors
-
Computational Thinking: A Meta-Review of Systematic Reviews and Meta-Analyses — Meta-review situating CT in CS education
-
Do Not Copy/Paste: Soft Barriers for Copying in AI-Assisted Programming — Copy-paste resistance in AI-assisted programming
-
Predicting Student Attrition in Competitive Programming: A Large-Scale Study Integrating Survey Insights and Global Behavioral Logs — Predicting Student Attrition in Competitive Programming
-
A Machine Learning Approach for Predicting Student Progress in Online Programming Education
-
Practical Implementation Report on Introducing Spec-Driven Development Using AI Agents in Software Development PBL - SDD with AI agents in a software PBL course; throughput vs. comprehension
-
Beyond Immediate Resolution: Generative AI as an Informal Cognitive Tutor in Novice Programming Learning — GenAI as informal cognitive tutor in novice programming learning
-
Flawed but Memorable: Student Critical Reception of Interest-Personalized GenAI Analogies in Computing Education — Flawed but Memorable: Student Critical Reception of Interest-Personalized GenAI Analogies in Computing Education
-
The Socratic trap: Benchmarking the capacity of large language models to generate strategic misconceptions in computer science education — SocraticTrap-CS: benchmarking models' capacity to generate strategic misconceptions across the CS curriculum
-
Ethical implications of prompt injection in AI-mediated grading: An adversarial red-team evaluation — Prompt injection red-team of AI-mediated grading: hidden instructions that move the mark undetected
-
AlgoRAG: Retrieval-Augmented Generation for Theoretical Computer Science Education -- A Comprehensive Evaluation Framework for Algorithm Analysis and Complexity Theory — AlgoRAG: Retrieval-Augmented Generation for Theoretical Computer Science Education -- A Comprehensive Evaluation Framework for Algorithm Analysis and Complexity Theory
-
LLMs Unplugged: Teaching Resources for a ChatGPT World — LLMs Unplugged: Teaching Resources for a ChatGPT World
-
Instructional Governance by Design: A Framework for AI in Computing Education — Instructional Governance by Design: A Framework for AI in Computing Education
-
Judgment-Centred Software Engineering Education: A Post-Hype Review and Framework for AI-Augmented Learning — A Post-Hype Review and Framework for AI-Augmented Software Engineering Education
-
Where LLM Graders Succeed and Break: Evidence from Two Computer-Science Exams — Where LLM Graders Succeed and Break: Evidence from Two Computer-Science Exams