Explaining Too Much? Understanding How Large Language Model Reasoning Traces Influence Performance and Metacognition

Created: 2026-05-26 | Tags: llmmetacognitionstudent-experienceefficacy-studyover-reliance

Fernandes, Buschek, Tankelevitch, Kosch & Welsch (2026) โ€” University of Bayreuth, Microsoft Research.

๐Ÿ“„ Full text (arXiv)

This preregistered between-subjects study (N=559) provides the first rigorous evidence that llm reasoning traces โ€” increasingly common in AI interfaces โ€” do not improve performance and can actively impair it. More critically, they create a dangerous metacognitive blind spot: participants substantially overestimate their performance regardless of trace format.

Key Findings

Connection to AIED

These findings have profound implications for intelligent-tutoring and AI feedback systems. If students feel more confident after seeing AI reasoning but don't actually learn better, then simply exposing AI reasoning in educational interfaces may create an over-reliance trap. The paper's recommendation โ€” that calibration should be scaffolded by interactions that elicit users' own reasoning first โ€” directly aligns with self-regulated-learning principles and cognitive offloading research showing that AI use can reduce active engagement.

Contrast with Assessment Governance

While GenAI assessment governance focuses on when to allow AI in evaluation, this paper addresses how AI explanations affect learning โ€” suggesting that even well-designed AI transparency features can backfire without metacognitive scaffolding.

Related Pages

Citation

APA: Fernandes, D., Buschek, D., Tankelevitch, L., Kosch, T., & Welsch, R. (2026). Explaining too much? Understanding how large language model reasoning traces influence performance and metacognition. arXiv:2605.25856. cs.HC.