Research Article
LLM Pedagogical Behavior in AI Tutoring Interactions
Synthesis: Students increasingly use large language models as on-demand tutors for coursework and problem solving, yet little is known about the level of assistance these models actually provide in authentic learning interactions. Lee and colleagues operationalize this dimension as a five-level Scaffolding scale, validated against human annotations, that characterizes responses by the degree of direct assistance they offer. Applied to 14,637 Large Language Models (LLMs) responses from 203 students in a university AI course, responses are overwhelmingly concentrated at high levels of assistance, with more than 95% classified as either Explaining or Solving. Scaffolding level is systematically associated with students' subsequent conversational behavior but provides little additional predictive information about exam performance beyond prior achievement and dialogue behavior, offering an empirical baseline for evaluating alternative tutoring designs.
Key Findings
- A five-level scaffolding scale was developed and validated against human annotations to characterize how directly LLM responses help a student complete a task.
- In authentic student-LLM interactions, more than 95% of the 14,637 analyzed responses were classified at the highest assistance levels (Explaining or Solving).
- Scaffolding level is systematically associated with students' subsequent conversational behavior in the tutoring dialogue.
- Scaffolding level provides little additional predictive information about performance on three subsequent exams beyond prior achievement and dialogue behavior.
- The scale offers a measurement framework for evaluating how alternative tutoring designs change the assistance LLMs provide.
Connected Concepts
- Intelligent Tutoring
- Scaffolding
- Large Language Models (LLMs)
- Generative AI
- Student-AI Interaction
- Teaching
- Metacognition
Connected Articles
- The Safety Gap: Restoring Productive Struggle Through Pedagogically Aligned Generative AI — The Safety Gap: Restoring Productive Struggle Through Pedagogically Aligned Generative AI
- The Tutoring Effectiveness Index: Predicting LLM Math Tutor Quality from Four Conversation Signals — The Tutoring Effectiveness Index
- Access is Not Enough: Human Support Improves Engagement with AI Tutoring — Access is Not Enough: Human Support Improves Engagement with AI Tutoring
- CSTutorBench: Benchmarking Small Language Models as Tutors for Block-Based Programming — CSTutorBench: Benchmarking Small Language Models as Tutors
- ChatGPT-generated help produces learning gains equivalent to human tutor-authored help on mathematics skills — ChatGPT-generated help produces learning gains equivalent to human tutor-authored help
Citation
Lee, Baek, Park, & Shin (2026). LLM Pedagogical Behavior in AI Tutoring Interactions.