Fernandes, Buschek, Tankelevitch, Kosch & Welsch (2026) โ University of Bayreuth, Microsoft Research.
๐ Full text (arXiv)
This preregistered between-subjects study (N=559) provides the first rigorous evidence that llm reasoning traces โ increasingly common in AI interfaces โ do not improve performance and can actively impair it. More critically, they create a dangerous metacognitive blind spot: participants substantially overestimate their performance regardless of trace format.
Key Findings
- Summary traces preserved task performance at the no-trace baseline while elevating trust and hedonic appeal โ changing how users feel without helping them perform.
- Full traces from a verbose open-weight model actually impaired performance relative to answer-only baselines.
- No trace format supported calibrated self-evaluation โ metacognitive overestimation was universal.
- Hedonic appeal, not trust, carried the indirect path to overestimation, consistent with a processing-fluency account: the pleasant experience of reading traces inflates confidence without improving understanding.
Connection to AIED
These findings have profound implications for intelligent-tutoring and AI feedback systems. If students feel more confident after seeing AI reasoning but don't actually learn better, then simply exposing AI reasoning in educational interfaces may create an over-reliance trap. The paper's recommendation โ that calibration should be scaffolded by interactions that elicit users' own reasoning first โ directly aligns with self-regulated-learning principles and cognitive offloading research showing that AI use can reduce active engagement.Contrast with Assessment Governance
While GenAI assessment governance focuses on when to allow AI in evaluation, this paper addresses how AI explanations affect learning โ suggesting that even well-designed AI transparency features can backfire without metacognitive scaffolding.Related Pages
- metacognition โ Metacognition in learning
- over-reliance โ AI overreliance patterns
- self-regulated-learning โ Self-regulated learning
- cognitive-offloading-speedup-illusion โ Cognitive offloading and miscalibration
- llm-fallacy-misattribution โ LLM fallacy misattribution
- intelligent-tutoring โ Intelligent tutoring systems
- genai-assessment-governance โ GenAI assessment governance