๐Ÿง  AI Ed Wiki

Alexandra Neagu, Jeffrey T. H. Wong, Marcus Messer, Rhodri Nelson, Peter B. Johnson (2026). Pluralistic Alignment Workshop @ ICML 2026

Key Findings

  • Alignment and evaluation methods for embedding scaffolding behaviour into chatbots rest on an implicit assumption: that students will take up the scaffolding and engage in the conversation.
  • The paper introduces an evaluation pipeline around two metrics โ€” Chatbot Scaffolding and Student Uptake โ€” applied across nine datasets of 9,490 chats spanning AI tutor benchmarks and real-world deployments of educational chatbots.
  • While benchmarks assume a high-scaffolding, high-student-uptake environment, students in real-world settings exhibit lower levels of uptake overall, frequently bypassing the chatbot's pedagogical framing to drive the interaction toward their own learning goals at little interpersonal cost.
  • Bypassing scaffolding is not necessarily detrimental; it frequently highlights a mismatch between a chatbot's pedagogical framing and the student's learning goals.
  • Future benchmarks must move beyond the assumption that students will simply take up the scaffolding, and instead evaluate how chatbots navigate diverse learning contexts and student-driven interaction patterns.
  • Study Design & Method

    Scaffolding describes how a tutor calibrates support to the learner's current state โ€” guiding through graduated hints, posing questions rather than giving answers, and withdrawing support as the student gains competence. Delivering timely, dialogic, and scaffolded feedback to every student at every moment of struggle is difficult at scale, and LLM-based chatbots have been proposed as a way to approach this challenge. However, deploying LLMs as tutors introduces a tension: they are trained to be helpful by presenting information and answering directly, rather than engaging students in guided discovery โ€” behaviour that is at odds with scaffolding, where a tutor withholds answers to promote reasoning. The evaluation pipeline operationalizes this tension through the Chatbot Scaffolding and Student Uptake metrics, and the corpus spans both benchmark datasets and real-world chatbot deployments.

    Relevance to AI in Education

    This paper contributes directly to understanding how AI systems interact with learners in authentic educational settings. It challenges benchmark assumptions about student uptake of LLM tutor scaffolding, showing that real-world learners frequently bypass pedagogical framing in favour of their own goals, and that this behaviour is often a rational response to a mismatch rather than a failure of engagement. For AI Tutoring design, the implication is that scaffolding should be adaptive to student-driven interaction patterns โ€” including Help Seeking styles โ€” rather than presupposed by the interface. The conversational structure of tutoring normally allows students to respond, negotiate, and ask follow-up questions, building understanding iteratively and exercising agency; benchmarks that ignore this dynamic risk overestimating both the value of rigid scaffolding and the quality of LLM tutors. For the Benchmark community, the work argues for evaluation designs that reward chatbots for navigating diverse learning contexts instead of assuming uptake.

    Connected Concepts

  • Help Seeking
  • AI Tutoring
  • Pedagogical LLM Training
  • Benchmark
  • Affective Tutoring
  • Socratic AI Dialogue
  • Pedagogical Agent
  • Automated Question Generation
  • Connected Articles

  • LLM Judged Helpfulness Pedagogy Signal โ€” Rethinking LLM-Judged Helpfulness as a Pedagogy Signal: A Pre-Registered Audit Across Tutor Models
  • Measuring LLM Tutors Teach Vs Solve โ€” Measuring Whether LLM Tutors Teach or Solve: A Diagnostic for Educational Impact
  • Student Misconceptions Conditionals Loops Taxonomy โ€” How Students (Mis)understand Conditionals and Loops -- A Taxonomy
  • Multi Agent LLM Social Learning โ€” Beyond the AI Tutor: Social Learning with LLM Agents
  • Zhang Tutormoments 2026 โ€” When Help is Unhelpful: Evaluating AI Tutors for Productive Struggle
  • Didactical Teacher Assistant Dimensional Modeling โ€” A didactical-driven teacher assistant for a dimensional modeling course
  • Citation

    Alexandra Neagu, Jeffrey T. H. Wong, Marcus Messer, Rhodri Nelson, Peter B. Johnson (2026). Rethinking Scaffolding in LLM Tutors: The Interactional Mismatch Between Benchmarks and Real-World Deployments. arXiv:2606.15766. Pluralistic Alignment Workshop @ ICML 2026.