📄 Research Article
Exploring Conversational Agents for Novice Programmers: A Scoping Review
Synthesis: This PRISMA-ScR scoping review (23 studies, 2019–June 2024) maps conversational agents for novice programmers, finding the field shifting from rule-based chatbots toward LLM- and RAG-based agents — yet only 4 of 23 studies ground design in learning theory. It flags a major inclusivity gap (17 of 23 prototypes English-only despite most research originating outside English-speaking countries) and weak, non-standardized evaluation, and offers design recommendations for introductory programming education.
Key Findings
- 23 studies, PRISMA-ScR. The review screened 743 citations to select 23 studies on educational conversational agents for novice programmers (January 2019–June 2024), with research peaking in 2022.
- A technological shift toward LLMs and RAG. Prototypes moved from rule-based/scripted systems (n=9) to LLM-based (n=8) and retrieval-augmented/hybrid (n=2) architectures — integrating open-source models, retrieval-augmented generation to reduce hallucination, and flexible pipelines that align technical sophistication with pedagogical adaptability.
- Sparse pedagogical grounding. Only 4 of 23 studies explicitly apply learning theories (Vygotskian dialogue, the 4C/ID model, Scaffolding techniques, gamification). Without theory-informed design, these tools risk bypassing critical and algorithmic thinking development — surfacing a persistent gap between CA development and Pedagogy.
- Weak and heterogeneous evaluation. 15 of 23 studies used experimental designs, but most relied on subjective post-usage surveys/interviews; only three used objective pre/post-test designs. Quality was mostly Moderate (none of 17 quasi-experimental studies rated High), limiting causal claims.
- English-only dominance creates an inclusivity barrier. Though most studies originated outside English-speaking countries, 17 of 23 prototypes were English-only; multilingual support (Pynar, Pyo, Profe Alex) remains rare, and only 2 of 23 studies addressed gender representation.
Implications
For introductory programming education, this review shows conversational agents are increasingly viable as personalized tutors offering Feedback and adaptive guidance, but their effectiveness hinges on grounding in pedagogical theory rather than technical novelty alone. The recommendation to integrate instructional-design principles, cognitive-load management, and formative feedback into modular, educator-customizable templates connects directly to the wiki's Pedagogical Agent and Conversational AI threads.
The inclusivity findings reinforce that language and gender equity must be designed-in — multilingual support and gender-neutral, inclusive interaction — rather than treated as afterthoughts. The call for standardized evaluation frameworks (usability + long-term learning outcomes + cognitive engagement) and for interdisciplinary collaboration (education, CS, HCI, psychology) aligns with AI Ed Evaluation concerns across the wiki. The "vibe coding" collaborative pattern connects to vibe coding work.
Connected Concepts
- Conversational AI
- Pedagogical Agent
- Intelligent Tutoring
- CS Education
- Scaffolding
- Feedback
- Cognitive Offloading
- Self Regulated Learning
- Metacognition
- Generative AI
- LLM
- RAG
- Multimodal
- Equity In AI Education
- AI Ed Evaluation
Connected Articles
- Gaide Vibe Coding K12 Teachers — Vibe coding framework for K-12 teachers
- Conversational AI Tutors Framework — Conversational AI tutors framework
- Tutoring Specific Vs General AI — Tutoring-specific vs general AI
- Measuring LLM Tutors Teach Vs Solve — Measuring whether LLM tutors teach or solve
- Conversational AI Agents Umbrella Review 2026 — Umbrella review of conversational AI agents in education
Citation
Barzanji, C., & Loitsch, C. (2025). Exploring conversational agents for novice programmers: a scoping review. Discover Artificial Intelligence, 5, 271.