Research Article
Open Questions Towards Skill-Sustaining Reliance in Reflective AI Engagement
Synthesis: Cognitive forcing, pre-commitment, and uncertainty communication reduce overreliance within a single session, but Sander de Jong's position paper argues the field has never checked whether they sustain expertise over time. It raises four open questions about Scaffolding that withholds or prompts independent reasoning, the Learner Agency it takes from experienced professionals, the pressure to prioritize throughput over friction, and the risk that repeated prompts habituate users instead of building durable Metacognition. The worry is that a mechanism supplying reflection replaces the self-monitoring it was meant to protect, so short-term gains in decision quality may coexist with long-term skill erosion.
Key Findings
- Four open questions, no new data. A workshop position piece on human agency and skill in AI-supported work; it turns prior findings into programmatic questions, not a study.
- Reflective support works unevenly. Cognitive forcing reduced overreliance on average, but benefited mainly participants higher in Need for Cognition, and engagement shifted with task difficulty.
- Adaptation trades against autonomy. Adaptive scaffolding requires the system to assess user competence, which the paper says reduces professionals' authority over their own development and risks over-scaffolding experts.
- Friction loses in throughput-driven work. Reflective mechanisms slow decisions deliberately; where performance is measured by speed, and rivals offer the same support without friction, skill-sustaining designs become a competitive disadvantage.
- Prompts can replace the capacity they protect. Repeated prompts wear off through habituation, so a professional who reflects only when the system asks may end with neither the tool nor self-monitoring habits.
- Reliance measures miss the outcome that matters. Longitudinal studies track whether users rely appropriately over time, not whether their capacity for independent reasoning survives.
- The evidence asymmetry is the core problem. Deployers decide now on session-level metrics — task accuracy, reliance rate, trust calibration — while the long-term skill case stays untested.
Promising effects measured too early
Pre-commitment, uncertainty communication, and added reflection steps such as cognitive forcing mitigate overreliance and improve decision quality, but they demand extra effort from workers; the problem is whether a mechanism survives a real workflow, not only whether it works. Rationale explanations generated by an Large Language Models (LLMs) have recently been used to prompt reflection, yet their fluency can make incorrect outputs look credible and discourage the critical engagement they were meant to provoke. What the field has skipped, the author argues, is whether professionals relying on reflective support still develop their own expertise or slowly lose the capacity for independent judgment.
Autonomy, adaptation, and who decides
Reflective mechanisms carry an implicit assumption that the user needs help thinking — potentially unfounded, even undermining, for an experienced professional. A novice still building foundational knowledge may profit from a system that withholds AI suggestions, while an expert able to evaluate the same advice critically may find withholding merely tedious. The obvious design answer, adaptive scaffolding, cuts against Learner Agency, because adaptation requires the system to assess user competence, and a mechanism perceived as unnecessary will simply be ignored. The open question is who decides which users receive scaffolding — the system, the organization, or the professionals.
Friction against throughput and market pressure
AI adoption in professional work is driven by faster decisions, higher throughput, and lower cognitive load. Reflective interaction contradicts that promise: it introduces friction, asks users to slow down, and demands effort at the moment when the AI could have answered immediately. A controlled study can absorb that trade-off; an environment measuring performance by quantity or speed cannot, and when reflection adds time to many daily decisions the cost is hard to justify. The tension is sharpest in high-stakes, risk-averse domains such as healthcare, where reflection and skill retention are tied to patient safety. If rival tools offer equivalent support prioritizing speed, skill-sustaining designs become a competitive disadvantage, raising the question of whether the remedy belongs to interaction design, organizational policy, or government.
Habituation and the short-term evidence gap
Scaffolding is meant to keep users engaged, but the prompts can become the source of engagement. Models rarely express uncertainty yet impose substantial metacognitive demands, so users lacking linguistic cues of doubt may fall into unwarranted reliance, and interaction often stops at a single prompt without follow-up. Repeated prompts may also wear off through habituation, much as repeated warnings are dismissed automatically — so scaffolding may replace the Metacognition it was meant to protect, leaving professionals without AI support or self-monitoring habits once the tool is removed. The author asks whether skill-sustaining reliance can itself become a form of deskilling, then notes that longitudinal work tracks trust and appropriate reliance over time but not whether independent reasoning survives. Session-level evidence against long-term claims leaves open what should justify deploying reflective mechanisms in professional practice, or deciding against them.
What this means for practice
- Designers. Treat reflective prompts as a fading skill-building intervention, not a permanent fixture, and test whether users reflect without the prompt.
- Instructors and trainers. Assess unaided reasoning alongside AI-supported work, so learners who perform well with assistance are not mistaken for those who retained it.
- Administrators. Price reflection into throughput targets rather than leaving it as unpaid extra effort; where only speed is measured, it will be abandoned.
- Researchers. Use longitudinal designs that measure skill rather than reliance, tracking whether independent reasoning holds up over repeated interactions, not just post-exposure accuracy and Trust Calibration.
- Developers. Communicate uncertainty plainly instead of shipping fluent rationales that manufacture credibility and suppress the engagement the feature was meant to encourage.
Limitations
- Workshop position piece with no empirical study of its own; the four open questions remain programmatic and unanswered.
- Its evidence comes from domains other than education — clinical decision-making, legal judgment, logistics, chess — so transfer to teaching is asserted, not demonstrated.
- The claim that prompts substitute for metacognitive habits rests on short-term habituation and overconfidence findings; no longitudinal scaffolding data is presented.
- Individual differences appear mainly as Need for Cognition and perceived task difficulty, leaving other moderators untested.
Connected Concepts
- Learner Agency
- Metacognition
- Scaffolding
- Cognitive Offloading
- Cognitive Surrender
- Human AI Collaboration
- Human-in-the-Loop
- Trust Calibration
- Trust
- Large Language Models (LLMs)
- Workplace Learning
- Self-Regulated Learning
- Critical Thinking
- Evaluative Judgment
- Desirable Difficulties
Connected Articles
- AI Advice Suppresses People's Willingness to Say "I Don't Know", Even When the Advice Is Wrong and Accuracy Is Incentivized — AI Advice Suppresses People's Willingness to Say "I Don't Know", Even When the Advice Is Wrong and Accuracy Is Incentivized
- Designing Against Deskilling: Metacognitive Feedback Reduces Cognitive Offloading to LLM Assistants — Designing Against Deskilling: Metacognitive Feedback Reduces Cognitive Offloading to LLM Assistants
- PAUSE: A Privacy-Preserving Self-Reflection Tool for AI-Associated Cognitive Offloading — PAUSE: A Privacy-Preserving Self-Reflection Tool for AI-Associated Cognitive Offloading
- Beyond checking: verification quality, reliance calibration, and learning in generative AI-assisted higher education — Beyond checking: verification quality, reliance calibration, and learning in generative AI-assisted higher education
- Benefit or Bottleneck? Assessing the Impact of Structured Reflection on Learning from AI-Driven Explanatory Feedback — Benefit or Bottleneck? Assessing the Impact of Structured Reflection on Learning from AI-Driven Explanatory Feedback
- Against frictionless AI — Against frictionless AI
- Trust and Reliance on AI in Education: AI Literacy and Need for Cognition as Moderators — Trust and Reliance on AI in Education: AI Literacy and Need for Cognition as Moderators
Citation
de Jong, S. (2026). Open Questions Towards Skill-Sustaining Reliance in Reflective AI Engagement. arXiv preprint.