🧠 AI Ed Wiki

LLM Reasoning Traces & Metacognition

This preregistered between-subjects study (N=559) provides the first rigorous evidence that LLM reasoning traces — increasingly common in AI interfaces — do not improve performance and can actively impair it. More critically, they create a dangerous metacognitive blind spot: participants substantially overestimate their performance regardless of trace format.

Key Findings

  • Summary traces preserved task performance at the no-trace baseline while elevating trust and hedonic appeal — changing how users feel without helping them perform.
  • Full traces from a verbose open-weight model actually impaired performance relative to answer-only baselines.
  • No trace format supported calibrated self-evaluation — metacognitive overestimation was universal.
  • Hedonic appeal, not trust, carried the indirect path to overestimation, consistent with a processing-fluency account: the pleasant experience of reading traces inflates confidence without improving understanding.
  • Connection to AIED

    These findings have profound implications for Intelligent Tutoring and AI feedback systems. If students feel more confident after seeing AI reasoning but don't actually learn better, then simply exposing AI reasoning in educational interfaces may create an Over Reliance trap. The paper's recommendation — that calibration should be scaffolded by interactions that elicit users' own reasoning first — directly aligns with Self Regulated Learning principles and cognitive offloading research showing that AI use can reduce active engagement.

    Contrast with Assessment Governance

    While GenAI assessment governance focuses on when to allow AI in evaluation, this paper addresses how AI explanations affect learning — suggesting that even well-designed AI transparency features can backfire without metacognitive scaffolding.

    Connected Concepts

  • LLM
  • Metacognition
  • Intelligent Tutoring
  • Over Reliance
  • Self Regulated Learning
  • Connected Articles

  • AI Peer Feedback Systems
  • Cognitive Offloading Speedup Illusion
  • GenAI Assessment Governance
  • Citation

    Fernandes, D., Buschek, D., Tankelevitch, L., Kosch, T., & Welsch, R. (2026). Explaining too much? Understanding how large language model reasoning traces influence performance and metacognition. arXiv:2605.25856. cs.HC.