On this page

Conversational AI (CAI) agents — AI-driven speech- or text-based agents that simulate and automate conversations, from rule-based chatbots to NLP/ML and Multimodal AI LLM-based assistants — are among the most widely used AI interfaces in education, valued for teaching, psychological, and metacognitive support even as technical, cognitive, and ethical concerns persist.

Questions to Consider

  • When you've used a chatbot or AI assistant, did you think of it as a teacher, a search engine, or something else? How did that framing shape how much you actually learned from it?
  • Conversational AI describes HOW an agent talks, not what it's built to do. Could a chatbot be conversational yet completely unpedagogical — and what would tell you the difference?
  • One study found that AI literacy — not general tech savvy — predicted whether students were willing and able to use a chatbot. Why might knowing how AI works matter more than being 'good with computers'?
  • A raw general chatbot can short-circuit reasoning by answering immediately, while a structured tutor preserves productive struggle by withholding answers. What design choices decide which kind of agent a student meets?
  • Students using a law-course chatbot did a third of their interactions after hours — evidence that 24/7 availability is a real benefit. But does always-on access also carry risks you'd want to design against?
  • The biggest barrier to chatbot adoption in one study was a badly designed pop-up, not distrust or academic-integrity fears. What does that suggest about where AI-in-education investments actually fail?

Introduction

Conversational AI (CAI) is the umbrella term for AI-driven agents that carry on spoken or written dialogue, most commonly realized as chatbots and, more recently, generative Large Language Models (LLMs)-based assistants such as ChatGPT, Claude, and multimodal educational avatars. Modern CAI agents fall into Machine Learning-based, NLP-based, and hybrid categories, with text-based agents the most prevalent in education. As learning tools they function as intelligent tutors, Feedback providers, interaction partners, and administrative assistants — overlapping with pedagogical agents while spanning a broader set of applications.

How conversational AI appears in the knowledge base

An umbrella-review synthesis. The umbrella review of CAI agents (34 review articles) shows CAI utilization is concentrated in teaching and learning support (97.1% of reviews), psychological and motivational support (91.2%), and metacognitive and personal development (88.2%), while administrative support, research management, and healthcare education lag. The review documents that human–AI relationship concerns persist across all CAI generations, with Academic Integrity and data Privacy emerging as newer ethical issues, and calls for HCI-grounded, evidence-based design and stronger AI Literacy support.

From chatbots to tutoring agents. The knowledge base traces CAI's evolution from rule-based FAQ chatbots toward tutoring-focused agents. The conversational AI tutors framework argues proven ITS technologies (Knowledge Tracing, affect detection, student modeling) should anchor generative tutors while Generative AI supplies flexible dialogue. Research on whether LLM tutors teach or solve and tutoring-specific vs general AI shows pedagogically designed Guardrails matter: raw general chatbots can short-circuit reasoning while structured tutors preserve productive struggle.

Benchmarks may overstate how much scaffolding gets taken up: across 9,490 chats in nine datasets, students in real deployments bypassed the chatbot's pedagogical framing at little interpersonal cost to drive the interaction toward their own goals — an instrumental Help-Seeking move rather than disengagement (Neagu et al. (2026)).

Interaction and collaboration. Conversational agents are increasingly framed as interaction partners rather than answer-givers. Student-AI Interaction captures how learners prompt, question, and verify with CAI in practice. In Collaborative Learning, agents mediate participation and shared AI Regulation in Education, and in Language Learning they provide real-time conversational practice. The Human AI Collaboration thread examines when this partnership preserves versus substitutes for the learner's cognitive work. LLMs as critique partners illustrate the preservation side concretely: Oppenheimer, Cash & Connell Pensky (2025) had ChatGPT, Gemini, or Claude critique students' argumentative essays across a semester, and learners improved on writing, prompt engineering, and response-to-feedback while rating the exchanges useful, engaging, and enjoyable — with active rebuttal of model claims (87.8%) showing they treated the conversational partner critically rather than passively.

Conversational agents are used to grade, and their evaluative behavior diverges from humans': ChatGPT was the most lenient of three evaluators of the same 52 undergraduate projects (M = 91.46 against 83.13 for the instructor) and shared no grading logic with peers (r = 0.05), though students valued its dialogic interactivity (Usher & Faraon (2026)).

Scaffolding the student's side of the conversation. Jin et al. (2026) add a question-specificity monitor and a revision template to a chat interface and intervene only on vague questions: in a between-subjects study with 40 college students learning web programming, HelpCoach users wrote significantly more specific first questions (57.3% vs. 40.5%) and retained more knowledge one week later (mean gain 1.9 vs. 0.3 of 6, d = 1.100). The retention benefit did not run through more targeted chatbot replies, which did not differ by condition — leaving the mechanism an open hypothesis rather than a demonstrated pathway.

The role the agent plays shapes the interaction. CAI design is not neutral about its persona: Liao (2026) found a fixed "student peer" companion in elementary book talk sustained longer interactions yet dominated the conversation (lower student word/sentence share) and hit an "affective ceiling" — matching a human teacher on factual recall but falling short on emotional and future-oriented reflection — arguing CAI should adapt its role (peer, teacher assistant, parent advisor) rather than stay monolithic. Xu et al. extend this to small groups, showing generative AI acts as both an agent and a collaborative space in synchronous and asynchronous collaborative dynamics, where the interaction design decides whether it scaffolds or supplants group cognition. Tone can also be set by an external classifier rather than by the conversation alone: Sukoon (Bashir & Afzal, 2026) trains a Random Forest on 20 survey features to assign a low, moderate, or high stress level, and that output selects one of three response tiers inspired by the Stepped Care Model — warm and encouraging at low stress, grounding and non-judgmental at high stress — before an Open Source LLM (GLM-4.5-Air via OpenRouter) takes over the dialogue with the full conversation history and a culturally adapted system prompt re-sent on every turn. The authors chose a free-access multilingual model deliberately to keep deployment feasible in regional universities, and note that Urdu appears through prompt engineering rather than a genuinely bilingual pipeline — a limitation that matters whenever cultural appropriateness is claimed as a system property rather than an evaluated outcome.

Conversational behavior tracks the provider, not the tuning. Across 1,971 AI and 135 human transcripts scored on six fixed rules drawn from the learning sciences, models from the same provider clustered together with no overlap between family ranges, which the authors read as teaching behavior following provider-wide training practices rather than individual model tuning (Northcutt et al. (2026)).

Response consistency is an infrastructure problem, not a prompting one: within-model reply similarity ran 0.715–0.795 against 0.443–0.604 between models, and adding chat history changed reply content even for the same focal message (Hao (2026)). Student perspectives. Real-world usage shows adoption hinges on AI Literacy and user experience more than on technical capability. A human-centered mixed-methods study of the "Jordan Chatbot," a GPT-4o-based pedagogical agent in an Australian law course, found students hold positive attitudes and perceive gains in knowledge while strongly supporting Academic Integrity requirements; over a third of interactions occurred after hours, confirming the value of 24/7 availability (Colbran, Jha & Schiavone 2026). Notably, AI literacy — not general technology proficiency — predicted willingness and confidence to use the chatbot, and usability (an intrusive pop-up design) was the largest barrier among non-users, ahead of trust, preference for staff, and academic-integrity fears.(Understanding student perspectives on generative AI chatbots: a human-centred mixed-methods study in higher education) The study recommends human-centered design, explicit AI policies and assessment labels, staff and student training, and continuous error monitoring — evidence that effective CAI deployment is as much a design and literacy problem as a technical one. At the other end of the age spectrum, Vahedian Movahed & Martin (2025) studied children aged 6–14 interacting with AMA, a topic-bounded, age-tailored chatbot (astronomy, sneakers and shoes, dinosaurs), finding broad openness to and high trust in the AI as an information source — children even tested its credibility with known-answer questions — alongside gaps in critical engagement and digital-safety awareness that argue for age-sensitive, trust-aware conversational-ai design and explicit Privacy instruction.

Who uses the agent also varies: among 97 graduate statistics students, more than a quarter never used the StatBot chatbot, and female students used it significantly more often (W = 679.5, p = .015), which the authors read as conversational AI lowering the perceived social risk of asking questions (Lee and Wu (2026)).

Adoption is not homogeneous: clustering 192 student and educator responses yields four personas — Cautious Achievers, Skeptical Utilitarians, Disengaged Doubters and Engaged Enthusiasts — that average-effect acceptance models obscure, with students prioritizing immediate Feedback and educators content accuracy and Academic Integrity (Saihi & Ahmed (2026)). Durability has been observed for one institutional deployment. A four-year randomized evaluation of a non-generative knowledge-base chatbot at California State University, Northridge (N = 8,708), found students remained receptive across eight semesters — annual opt-out rates never exceeded 4% — while impacts concentrated in time-sensitive administrative tasks such as early course registration (34 percentage points more likely to enroll by the deadline) and no statistically significant effects appeared on GPA, units, or persistence (Mata, Russell & Page 2026, a working paper). Its measurement lesson bears on how CAI engagement is counted: with opt-out near zero alongside roughly 5% active engagement on interactive campaigns, conclusions about sustained engagement depend on whether a study counts only direct replies or also passive engagement, where students act on information without texting back.

Risks and ethics. CAI agents carry persistent risks of over-reliance and cognitive offloading (the leading ethical concern in the umbrella review), plus technical limitations, hallucination, bias, plagiarism, and equity barriers. These concerns animate AI Literacy and Reducing AI Misuse and require policy and ethical-AI Governance responses. A further trust concern arises around the adoption advice staff increasingly receive: staff are urged to consult conversational AI about whether to adopt AI, yet such systems are built by organizations with a commercial stake in adoption. An audit of ten frontier LLMs found most acknowledge skeptical users' concerns before redirecting to engagement framings, raising questions about the neutrality of AI adoption advice.

Game-based conversational agents. Beyond tutoring, conversational agents are being embedded in digital game-based learning. Wenzel, Geiger, and Liening (2026) use action design research to derive the CAIS-GBL framework — four design principles and fifteen design features for AI conversational agents in digital game-based learning — grounded in theory-driven meta-requirements spanning cognitive, motivational, affective, and socio-cultural engagement and an equity-by-design stance. Their instantiated agent (Lara) in a business Simulation game was positively received for cognitive and social presence and support for self-regulated learning, evaluated with student teachers and in a field study — a practical blueprint for adaptive instructional support via conversational agents in serious games.

Relationship to pedagogical agents and intelligent tutoring

Conversational AI is best understood as an interaction modality that overlaps — but does not coincide with — two more established constructs in the knowledge base: pedagogical agents and intelligent tutoring systems (ITS).

Conversational AI as the medium, not the pedagogy. CAI names how the agent communicates (natural-language dialogue, spoken or text). It says little on its own about what the agent is built to do. Pedagogical agents, by contrast, are defined by their instructional role — an AI component that engages learners through dialogue, questions, or prompts to support metacognitive processes, Feedback, and Scaffolding. Intelligent tutoring systems are defined by their architecture and modeling — a diagnostic backbone of Knowledge Tracing, student modeling, and pedagogical decision logic that tracks what the learner knows and adapts instruction. A single agent can be all three at once: e.g. a conversational AI tutor is a CAI agent (dialogue interface) that functions as a pedagogical agent (tutoring strategies) built on an ITS foundation (student modeling). The distinction matters because a CAI agent need not be pedagogically grounded at all — a plain FAQ chatbot is conversational AI without being a pedagogical agent or a tutor.

The pedagogical-agent lens. Pedagogical agents use the conversational medium to enact teaching strategies — eliciting self-assessments, Socratic questioning, role-specialized facilitation in multi-agent designs (Teacher, Assistant, Classmate, Analyzer). Not every CAI agent is a pedagogical agent, but the two heavily overlap: the umbrella review of CAI agents found teaching and learning support (97.1%) and metacognitive development (88.2%) dominate CAI applications, meaning most education-focused CAI agents function pedagogically. The novice-programmer scoping review sharpens this: only 4 of 23 conversational agents explicitly grounded design in learning theory — most were pedagogical in intent but not in foundation.

The ITS lens. Intelligent tutoring contributes the cognitive diagnostic machinery that raw conversational models lack. The conversational AI tutors framework argues proven ITS technologies should anchor generative tutors: knowledge tracing, affect detection, and student modeling supply the structure, while Generative AI and LLMs supply flexible dialogue. This is the key design tension — conversational AI provides natural, scalable interaction, but without ITS-style structure it risks over-scaffolding, hallucination, or bypassing the learner's productive struggle. Research such as Measuring Whether LLM Tutors Teach or Solve: A Diagnostic for Educational Impact and The Evidence Base on AI in K-12: A 2026 Review shows that pedagogy-oriented criteria (guiding questions, calibrated hints) must be designed in explicitly.

In short: conversational AI is the interface/medium, pedagogical agents are the role, and intelligent tutoring is the underlying modeling and instructional logic. Educationally valuable CAI agents sit at the intersection of all three — conversational in interface, pedagogical in intent, and tutor-like in their modeling of the learner.

Practical guidance

Choose conversational agents to support teaching, Motivation, and Metacognition rather than merely to answer questions, and design for HCI-grounded, participatory, user-centered interaction. Guard against over-reliance by pairing CAI with AI Literacy instruction and Feedback that keeps the learner cognitively productive. Attend to AI literacy and usability explicitly — since these — not general digital skill — drive adoption and non-use (Colbran, Jha & Schiavone 2026) — and pair deployment with clear AI-use policies, assessment labels, and training. Evaluate CAI on pedagogical outcomes — not just task completion — and plan for equity and Accessibility from the start rather than as an afterthought.

Connected Concepts

Connected Articles

Citation

Ganguly, A., Mehjabin, N., Malik, A., & Johri, A. (2025). Conversational AI agents in education: an umbrella review. AI and Ethics, 6, 72.

Embed this page

Copy the code below to embed a chromeless version of this page in a learning management system or other website. The embedded view hides the site header, navigation, and footer.