Research Article
Conversational AI agents in education: an umbrella review of current utilization, challenges, and future directions
Synthesis: This umbrella review (34 review articles, PRISMA 2020) maps how conversational AI (CAI) agents are used across education — concentrated in teaching, psychological, and metacognitive support while administrative and research functions lag — and documents that human–AI relationship concerns persist across all CAI generations while Academic Integrity and data Privacy emerge as newer ethical issues. It surfaces gaps in CAI frameworks (no end-to-end design guidance, weak CAI-specific usability, unclear classroom implementation) and proposes an ethical-and-responsible-use roadmap grounded in AI Literacy, participatory design, and continuous evaluation.
Key Findings
- A synthesis across a fragmented field. The umbrella review systematically integrates 34 prior review articles on conversational AI agents in education across five databases, bridging older rule-based chatbot research with the newer GenAI/LLM literature to form a cohesive state-of-the-art.
- Utilization is heavily pedagogical. Teaching and learning support (97.1%), psychological/motivational support (91.2%), and metacognitive and personal development (88.2%) dominate, while administrative support (50%), research and information management (52.9%), and healthcare/medical support (41.2%) remain underexplored — pointing to opportunities beyond instructional contexts.
- Technical and cognitive concerns dominate. The most-discussed challenges are technical limitations (97.1%), educational impact and cognitive concerns (91.2%), and interaction/usability challenges (85.3%). As agents shift to GenAI/LLM-based, emphasis moves toward hallucination, bias, privacy, and plagiarism — with Accessibility and cultural inclusivity staying secondary.
- Human–AI relationship is the persistent ethical thread. Human–AI relationship concerns (over-reliance, social isolation, depersonalization, transparency, accountability) are the most frequently discussed ethical concern across all CAI generations. Academic Integrity is discussed in only 35.3% of articles, signaling it as a comparatively newer concern.
- Framework gaps and a roadmap. Reviews lack end-to-end design guidance, CAI-specific usability methods, clear classroom implementation and Teaching strategies, and AI-literacy support. The proposed roadmap centers foundational AI literacy assessment, participatory (HCI-grounded) design, ethical-use guidelines, and continuous evaluation of cognitive impact.
What this means for practice
- Instructors. Use conversational agents where the evidence is strongest — teaching and learning support, psychological and motivational support, and metacognitive development — rather than for administrative or research tasks, which the 34 reviews mention far less often.
- Instructors. Deploy with an explicit plan for the human–AI relationship: over-reliance, social isolation, depersonalization, transparency, and accountability are the most frequently discussed ethical concerns across every CAI generation, so set expectations and monitoring for them before rollout.
- Designers. Design for classroom orchestration, not just conversation: the review finds no end-to-end design guidance and unclear Teaching and implementation strategies, so specify how the instructor's role changes and document the workflow.
- Researchers. Evaluate beyond short-term engagement: the roadmap calls for continuous assessment of cognitive impact, and the review flags long-term effects on critical thinking, self-regulated learning, and knowledge transfer as unresolved.
- Administrators. Adopt the four-pillar roadmap — AI Literacy assessment, participatory HCI-grounded design, ethical-use guidelines, and continuous evaluation — as a governance checklist when approving a conversational agent for a course or institution.
Limitations
- The evidence base is 34 review articles rather than primary studies, so findings inherit the limitations and selection criteria of those reviews and cannot establish effects not already synthesized.
- Reported figures are frequencies of mention across reviews (teaching support 97.1%, technical limitations 97.1%, academic integrity 35.3%), which measure how often a topic appears, not its magnitude or direction of effect.
- Coverage is bounded by the search across five databases and the eligibility criteria, and the review itself reports the field skews toward higher education and language learning, leaving K-12, STEM, and administrative contexts underrepresented.
- The proposed roadmap is not validated: the review offers it as guidance, and no implementation study tests whether its four pillars improve outcomes.
Citation
Ganguly, A., Mehjabin, N., Malik, A., & Johri, A. (2025). Conversational AI agents in education: an umbrella review of current utilization, challenges, and future directions for ethical and responsible use. AI and Ethics, 6, 72.