Research Article
AI chatbot design principles to enhance the collective efficacy in collaborative learning
Synthesis: AI chatbot design principles to enhance the collective efficacy in collaborative learning
Key Findings
- A design and development research study (Method Type 2, Richey & Klein 2014) producing a validated framework for chatbots that enhance collective efficacy — the shared belief among team members in their collective ability to accomplish tasks (Bandura, 2000) — in collaborative learning: 4 design elements, 10 design principles, and 46 actionable sub-guidelines.
- The final four design elements are Group Cohesion Support, Affective Cohesion Support, Support for Collaborative Learning Activities, and Inducing Conversation, spanning principles of belongingness formation, establishing interdependence, creating a positive atmosphere, empathy formation, promoting sharing, enhancing collaborative Problem Solving, supporting social AI Regulation in Education, providing immediate Scaffolding, familiarity, and personification.
- The principles were derived from a Scopus literature review (social sciences, English; terms including "chatbot," "design," "collective efficacy," "collaborative learning tools") that collected 929 papers, narrowed to 116, and analyzed 73 for chatbot design elements, principles, and guidelines.
- Validation proceeded through three rounds of expert review with five experts (Ph.D.-level, 10–13 years of experience in collaborative learning and AI chatbots): round one showed low validity (M = 2.80, SD = 0.84; content validity index 0.60; inter-rater agreement 0.00), while rounds two and three reached near-perfect scores — round three scored 4/4 on all domains except explanatory power (M = 3.80, SD = 0.45), with CVI and inter-rater agreement of 1.00.
- A usability test with eight participants (three AI chatbot designers and five instructors who design collaborative learning sessions, ~30-minute interviews each) confirmed strengths — learners addressing challenges independently, chatbot-mediated team facilitation, and precise guidance preventing unproductive time — while identifying limited feedback for non-participating learners as a weakness.
- The framework is designed to be usable by educators without technical expertise, supporting scalable, accessible implementation in real classrooms.
Study Design & Method
The study used Design and Development Research Method Type 2, proceeding through two stages: design principle development and validation. In the development stage, initial design principles were derived from a systematic literature review (929 Scopus papers collected; 116 selected after excluding technically focused papers; 73 analyzed after excluding non-educational-chatbot or non-collaborative-learning contexts), from which design elements, principles, and detailed guidelines were extracted and consolidated. The initial framework comprised three design elements (learning support, emotional support, rapport building), ten principles, and 54 sub-guidelines. In the validation stage, internal validity was tested through three rounds of expert validation using surveys adapted from Nail and Jung (2001) — rating validity, explanatory power, usefulness, and generalizability on a 4-point Likert scale — plus ~30-minute semi-structured interviews; external validity was tested through the usability test with designers and instructors, analyzed thematically and folded back into the final principles.
Key Results
- Iterative refinement: round one of expert validation forced reorganization of principles under design elements and consolidation of redundant sub-guidelines (e.g., Sense of Belonging merged into Interdependence); round two split and introduced principles (Sense of Belonging Formation, Social Regulation Support) and separated compound sub-guidelines; round three restructured the design elements — splitting collective-efficacy formation into group cohesion and collaborative problem-solving support and reclassifying immediate feedback under the dialogue-inducing element — yielding the final 4-element, 10-principle, 46-sub-guideline framework.
- Validation trajectory: validity rose from M = 2.80 (SD = 0.84) with CVI 0.60 and inter-rater agreement 0.00 in round one to 3.80 with CVI and inter-rater agreement of 1.00 in round two, and to 4/4 across domains (explanatory power M = 3.80, SD = 0.45) with perfect CVI and inter-rater agreement in round three.
- Usability findings: the chatbot design helped learners solve problems independently, mediated within-team discussion, and offered precise guidance that prevented unproductive time use; suggested improvements included personalized feedback for inactive learners.
- Interdisciplinary perspectives: educators emphasized pedagogical soundness and real classroom dynamics while developers and educational technologists emphasized usability, feasibility, and real-world implementation, jointly shaping the final principles.
What this means for practice
- Instructors. Give the chatbot the monitoring job a single instructor cannot do: have it read team conversation logs and deliver team-specific feedback, immediate Scaffolding, and social AI Regulation in Education support continuously rather than at the end of a session.
- Instructional designers. Treat affective cohesion as a design requirement equal to cognitive support, building belongingness formation, interdependence, a positive atmosphere, empathy formation, and personification into the four design elements rather than leaving them to chance.
- Designers. Give non-participating members their own prompt stream: the usability test's one flagged weakness was limited feedback for learners who were not contributing, so specify personalization that reaches inactive team members.
- Instructors. Adopt the framework as a specification you can implement without coding expertise — its 46 sub-guidelines are written for educators — while treating classroom effect as untested until a chatbot built from them is evaluated in a real course.
- Administrators. Settle informed consent for conversation-log data, algorithmic bias mitigation, and institutional data-governance policy before deployment, since the chatbot collects sensitive team discourse and the validation flagged culturally or emotionally inappropriate responses.
Limitations
- The design principles have not yet been implemented in and evaluated against real educational settings; the authors call for studies that build chatbots from the principles and test their effect on collaborative learning outcomes to assess generalizability.
- The principles are general rather than context-specific; project-based, maker, and discussion-based learning each have distinct characteristics that may require optimized or additional design principles.
- The usability test used a small, relatively homogeneous sample of eight participants, and expert validation involved a small expert group with limited diversity of perspectives, potentially biasing the results; broader stakeholder samples and iterative testing cycles are recommended.
Citation
Kim, M., & Lim, C. (2025). AI chatbot design principles to enhance the collective efficacy in collaborative learning.