On this page

Synthesis: This study tests whether a simple transparency intervention that warns students an AI pedagogical agent may make mistakes affects learner behavior in a math intelligent tutoring system. In a classroom experiment with 252 school students, those warned about potential AI errors requested significantly more hints than a control group, even though system behavior was identical — showing that lightweight transparency interventions can influence learners' interaction strategies, with implications for trust calibration and Help-Seeking in AI-supported learning.

Abstract

Recent work in Technology-Enhanced Learning and Human-Computer Interaction highlights the importance of transparency and trust calibration in AI-supported learning environments as they pose a risk of hallucinations. In this study, we investigate whether a simple transparency intervention that warns students that a pedagogical agent may make mistakes affects learner behavior in a math intelligent tutoring system. We conducted a classroom experiment with 252 school students using two system versions: one including a warning message about potential system errors, and one that does not mention potential errors. Using log data, we analyzed students' Problem Solving performance data, including help-seeking behavior, error rate, and time-on-task. Results show that students who were warned about potential AI errors requested significantly more hints than those in the other condition, even though the actual system behavior was exactly the same. This finding suggests that lightweight transparency interventions can influence learners' interaction strategies without necessarily improving or impairing immediate performance — a contribution to learner behavior and AI transparency in education.

What this means for practice

  • Instructors. Add a short fallibility warning to AI-supported activities even when the deployed system is error-free: hint requests rose significantly in the warned condition (β = −0.33, t(235) = −2.33, p = .02) with identical system behavior.
  • Instructors. Read the warning as a strategy lever, not a performance fix — error rate and time-on-task did not differ between conditions, so the intervention changed learner behavior rather than immediate outcomes.
  • Learners. Treat tutor hints as checkable guidance rather than authoritative answers; the study's design assumes a learner who reads the agent's messages thoughtfully and thinks critically about the guidance received.
  • Designers. Design the wording of uncertainty cues deliberately, since a single popup shown three times at the introduction of new solution methods was enough to shift interaction patterns.

Limitations

  • The final sample is 252 seventh-grade students (aged 12–13) in seven classes at one secondary school in Tokyo, taught by one teacher; 18 students were absent on data-collection days and excluded, from 270 originally recruited.
  • Exposure was two 50-minute class periods, and the authors state post-test data are not reported because unexpected class cancellations prevented many students from taking the post-test — so no learning claim is possible.
  • The tutor introduced no errors: hints and feedback were rule-based and hard-coded without LLMs, so the study cannot show how learners respond when the system actually makes mistakes.
  • Group sizes were unequal (12 students per class assigned to the warning condition) due to a concurrent data collection, which the authors note reduces statistical power.

Citation

Nagashima, T., Hladký, M., & Rief, V. (2026). Warning About AI Fallibility Increases Help-Seeking in an Intelligent Tutoring System.

Embed this page

Copy the code below to embed a chromeless version of this page in a learning management system or other website. The embedded view hides the site header, navigation, and footer.