AI Ed Wiki logoAI Ed WikiUse with AI

Synthesis: This PRISMA-ScR scoping review (23 studies, 2019–June 2024) maps conversational agents for novice programmers, finding the field shifting from rule-based chatbots toward LLM- and RAG-based agents — yet only 4 of 23 studies ground design in learning theory. It flags a major inclusivity gap (17 of 23 prototypes English-only despite most research originating outside English-speaking countries) and weak, non-standardized evaluation, and offers design recommendations for introductory programming education.

Key Findings

  1. 23 studies, PRISMA-ScR. The review screened 743 citations to select 23 studies on educational conversational agents for novice programmers (January 2019–June 2024), with research peaking in 2022.
  2. A technological shift toward LLMs and RAG. Prototypes moved from rule-based/scripted systems (n=9) to LLM-based (n=8) and retrieval-augmented/hybrid (n=2) architectures — integrating open-source models, retrieval-augmented generation to reduce hallucination, and flexible pipelines that align technical sophistication with pedagogical adaptability.
  3. Sparse pedagogical grounding. Only 4 of 23 studies explicitly apply learning theories (Vygotskian dialogue, the 4C/ID model, Scaffolding techniques, gamification). Without theory-informed design, these tools risk bypassing critical and algorithmic thinking development — surfacing a persistent gap between CA development and Pedagogy.
  4. Weak and heterogeneous evaluation. 15 of 23 studies used experimental designs, but most relied on subjective post-usage surveys/interviews; only three used objective pre/post-test designs. Quality was mostly Moderate (none of 17 quasi-experimental studies rated High), limiting causal claims.
  5. English-only dominance creates an inclusivity barrier. Though most studies originated outside English-speaking countries, 17 of 23 prototypes were English-only; multilingual support (Pynar, Pyo, Profe Alex) remains rare, and only 2 of 23 studies addressed gender representation.

Implications

For introductory programming education, this review shows conversational agents are increasingly viable as personalized tutors offering Feedback and adaptive guidance, but their effectiveness hinges on grounding in pedagogical theory rather than technical novelty alone. The recommendation to integrate instructional-design principles, cognitive-load management, and formative feedback into modular, educator-customizable templates connects directly to the wiki's Pedagogical Agent and Conversational AI threads.

The inclusivity findings reinforce that language and gender equity must be designed-in — multilingual support and gender-neutral, inclusive interaction — rather than treated as afterthoughts. The call for standardized evaluation frameworks (usability + long-term learning outcomes + cognitive engagement) and for interdisciplinary collaboration (education, CS, HCI, psychology) aligns with AI Ed Evaluation concerns across the wiki. The "vibe coding" collaborative pattern connects to vibe coding work.

Connected Concepts

Connected Articles

Citation

Barzanji, C., & Loitsch, C. (2025). Exploring conversational agents for novice programmers: a scoping review. Discover Artificial Intelligence, 5, 271.