Research Article
TeachingCoach: A Fine-Tuned Scaffolding Chatbot for Instructional Guidance to Instructors
Synthesis: Higher education instructors often lack timely and pedagogically grounded support. Universities operate teaching and learning centers that offer workshops and consultations, but these resources face key limitations: support may not be available at the moment of need, feedback is often generic rather than tailored to an instructor's background, and some educators hesitate to seek direct help for fear of appearing unskilled. General-purpose LLMs like ChatGPT are widely accessible, but their responses are typically generic and rarely apply evidence-based principles in a pedagogically scaffolded way. TeachingCoach is a pedagogically grounded chatbot that simulates the role of a teaching expert and delivers conversational guidance on instructional practice through problem identification, diagnosis, and strategy development.
Summary
The system is built on a data-centric pipeline. First, pedagogical rules are extracted from educational resources, with experts encoding 36 core instructional rules as structured system prompts that anchor conversations in evidence-based practices. Second, GPT-4o generates realistic and diverse training data: teacher profiles specifying years of experience and teaching subject, teaching challenges such as managing classroom attention, and multi-turn conversations between a teacher and a simulated expert that progress over 20–30 turns through clarifying the challenge, exploring strategies, planning next steps, and reflecting on practice. Human experts review the generated conversations, removing those that are inconsistent, repetitive, or pedagogically unsound; after filtering, the dataset contains 406,183 training, 4,156 validation, and 4,143 test examples. Finally, a LLaMA-2-13B-Chat model is fine-tuned with full parameter updates; at each turn the assistant first predicts an instructional step — Step 1: Identify the Problem, Step 2: Explore Reasons, Step 3: Develop Strategies — and then generates the response conditioned on that step, aligning dialogue progression with a structured pedagogical scaffold.
Key Contributions
- Outperforms a GPT-4o baseline on expert-rated dialogue quality: across 200 conversations rated on a 3-point scale over four dimensions (clarity, respectful tone, encouragement of reflection and reasoning, acknowledgment of user input), TeachingCoach consistently scored higher (e.g., 2.62 vs. 2.01 on clarity).
- A user study with higher education instructors highlights trade-offs between conversational depth and interaction efficiency: although the baseline model was preferred overall (21 vs. 13), participants more often attributed learning to the fine-tuned model (18 vs. 13), and agreement between preference and learning attribution was low (39%).
- Step labels generated by the model can be logged for analytics and personalization while optionally being hidden from the visible transcript, enabling supervision of the instructional process.
- The demo system provides onboarding (asking users about experience, current courses, and AI attitudes), the chat interface, and a dashboard for scheduling consultations with live experts, managing collected resources, and storing user data.
Study Design & Method
Expert evaluations compared TeachingCoach with a GPT-4o baseline in a zero-shot setting, highlighting the impact of explicit pedagogical supervision and step-aware training independent of model scale or proprietary data. The user study with higher education instructors examined both overall preference and perceived learning, finding that participants consistently distinguished the models on conversational engagement, efficiency of interaction, and breadth of suggestions: the fine-tuned model was described as resembling dialogue with a human pedagogy expert through reflective questions and follow-up prompts, while the baseline was valued for generating responses quickly when immediate guidance was needed. The work positions TeachingCoach against prior systems that primarily support students, arguing that instructor-facing support grounded in pedagogical practice has been underserved.
What this means for practice
- Instructors. Work a teaching problem through the structured sequence — identify the problem, explore reasons, develop strategies — rather than asking for an immediate list of tips, and expect the exchange to take more turns before it converges on concrete classroom practices.
- Instructors. Judge such a tool by what it teaches you, not by how fast or broad it feels: in the user study the baseline was preferred overall (21 vs. 13) while the fine-tuned model was more often credited with learning (18 vs. 13), and the two judgments agreed only 39% of the time.
- Faculty developers. Use the same three-step scaffold as the backbone of coaching conversations and of any AI support you deploy, since expert raters credited that supervision — not model scale — with the clarity advantage (2.62 vs. 2.01 on a 3-point scale).
- Faculty developers. Pair breadth with depth deliberately: participants valued the baseline for rapid, wide-ranging suggestions but felt overwhelmed by them, and valued the fine-tuned model for focused, coherent guidance through reflective questions.
- Administrators. Fund just-in-time, instructor-facing guidance to sit alongside teaching and learning centers, whose workshops and consultations are not available at the moment a teaching problem appears.
Limitations
- The user study recruited 41 higher education instructors through public postings and drew on a self-selected sample across U.S. research universities, community colleges, and liberal arts colleges; the authors note this may limit generalizability to K-12 settings, informal learning environments, or institutions with different cultural and technological infrastructures.
- The evaluation captured instructors' perceptions only — up to 10 minutes with each system, then a brief survey and a 15-minute interview — and did not directly assess student learning outcomes or classroom practice over time.
- Training data was fully synthetic, generated by GPT-4o from instructor profiles and teaching challenges (406,183 training examples after expert filtering), so the fine-tuned model inherits the generator's coverage rather than real teaching conversations.
- The only comparison is against a GPT-4o-mini baseline, so the advantage over other step-aware or expert-designed systems remains untested.
Citation
Isabel Molnar et al. (2026). TeachingCoach: A Fine-Tuned Scaffolding Chatbot for Instructional Guidance to Instructors.