On this page

Synthesis: The same Generative AI system can deepen or erode Critical Thinking, Esmaeiligoujar and Rahimi argue, because the effect is a consequence of design rather than a property of the technology. Their systematic scoping review screened 1,392 records against PRISMA and CASP criteria, retaining 38 empirical studies published between 2017 and 2025 and sorting them into three themes: 22 studies in which GenAI supported reasoning, 8 in which it undermined it, and 8 capturing stakeholder views. Gains appeared when tasks required students to evaluate, question, revise or argue against AI output, preserving Evaluative Judgment as the learner's own work; losses appeared when fluent output arrived before any reasoning demand, letting Cognitive Offloading displace inquiry and self-monitoring. From this the authors build a Vicious or Virtuous framework in which five categories of moderating factor set the conditions, design determines whether a Virtuous Cycle of Scaffolding, dialogue, adaptive feedback and reflection activates or a Vicious Cycle of cognitive substitution, over-directive interaction, misaligned feedback and uncritical acceptance takes hold, and short-term engagement patterns compound into durable dispositions that are hardest to change during the K-12 years.

Key Findings

  1. Design, not the tool, is decisive. Across 38 studies the review found the same GenAI system producing opposite outcomes on Critical Thinking depending on whether the interaction required students to question, evaluate and revise output or merely to accept it. The authors frame Theme 1 (support, n = 22) and Theme 2 (harm, n = 8) as mirrors, each mechanism of support paired with a corresponding mechanism of harm.
  2. Evaluative engagement with AI-generated steps produced measured reasoning gains. Tashtoush et al. (2025) ran a quasi-experiment with 91 Grade 11 mathematics students in Jordan using Wolfram Alpha and Microsoft Math Solver over a six-week calculus unit, requiring students to judge each solution step rather than copy it. Post-tests favored the GenAI group across all four measured dimensions — deduction (M = 15.66 vs. 9.95), interpretation (17.32 vs. 10.95), inference (16.09 vs. 9.58) and evaluation (15.88 vs. 9.14) — with effect sizes of η² = 0.21 to 0.35.
  3. Error detection built evaluative stances. Sivenas (2025) studied 109 sixteen-year-old students across three Greek high schools using ChatGPT-5 in an eight-hour intervention with tasks designed to elicit hallucinations; many students then adopted "epistemic safeguarding", restricting AI use to domains where they could verify answers. Chang et al. (2025) found iterative, criteria-based revision of ChatGPT drafts strengthened high school students' Self-Regulated Learning and understanding of argument structure.
  4. Dialogue, prompt refinement and structured feedback all preserved reasoning. Tang and Putra (2025) built the Dialogic Science Teacher chatbot on Bakhtin's heteroglossia and observed perspective-taking, reasoning, arguing and Creativity in 21 secondary students' chat logs across two Indonesian schools; Li et al. (2024) required K–12 students to refine prompts before the system would respond; Behnamnia et al. (2024) found 300 elementary and middle school students in Iran who ran a predict-observe-explain cycle on BrainPOP out-performed open-access peers, with teacher structuring as a significant moderator.
  5. Undermining followed from unstructured access. Stoyanova et al. (2024) found Bulgarian high school students accepted AI health explanations without cross-referencing; Daskalaki et al. (2024) reported the same pattern in a survey of 1,754 educators across five European countries; Velander et al. (2024) observed students' deference to AI strengthening across sessions at a Swedish K-12 school.
  6. Cognitive substitution was documented in varied tasks. Mwakalinga and Mabilika (2025) found a majority of Tanzanian secondary students submitted AI-generated work with minimal engagement; Han and Han (2025) saw AIStoryBot steer 12 Icelandic middle schoolers' narratives toward predictable endings they then accepted; Behboudi et al. (2024) found critical consciousness activated only when grades 5–10 students were required to evaluate AI framings of social justice issues against their own knowledge.
  7. Feedback design inverted formative assessment when unprioritised. In Yin et al.'s (2025) K–12 writing feedback program across US grades 7, 8, 10 and 11, dense, unprioritised Feedback led students to read without revising and showed less Evaluative Judgment than peers given targeted, prioritized feedback requiring a response.
  8. Student awareness lagged teacher awareness. Sok et al. (2025) surveyed 315 Cambodian high school students: concern about reduced critical thinking and creativity was the lowest-rated item (M = 3.12 of 5), below data privacy and over-reliance (both M = 3.27), while teachers in Tripathi et al. (2025) described students "just blindly copying without anything registering".
  9. Frameworks and institutions converged on requirements, not bans. Curi et al. (2025), from Uruguay's national K–12 AI initiative, specify three epistemic questions ("What is the AI? How does the AI work? What can the AI do?") and five curriculum dimensions; Saddhono et al. (2024) found AI Literacy a significant positive predictor of critical thinking in a PLS-SEM study of 368 Indonesian secondary students; Yang et al. (2025) found 153 Taiwanese high schoolers learned programming more effectively when they attempted problems before comparing with ChatGPT.

How the review was conducted

The study is a systematic scoping review following PRISMA reporting (Moher et al., 2010), searching Scopus, Web of Science, IEEE Xplore, ScienceDirect and Google Scholar for 2017–2025, the window opened by "Attention Is All You Need" (Vaswani et al., 2017). Of 1,392 records, 35 met inclusion criteria and a targeted hand search added 3, giving 38 studies appraised with the CASP checklist. Two researchers coded independently, compared decisions, and resolved disagreements to consensus, yielding the three themes. The design deliberately excluded higher-education, purely technical and non-critical-thinking studies, which the authors note makes the K–12 evidence base thinner than the higher-education literature and its theoretical grounding inconsistent, with critical thinking often bundled with Creativity or Problem Solving as an umbrella construct.

The review rests on established Research Methods in AIED rather than invention. The authors read findings through Salomon's (1990) distinction between cognitive effects with and of technology, Sweller's (1988) separation of germane from extraneous load, Vygotsky's (1978) zone of proximal development and Wood et al.'s (1976) fading scaffold, Flavell (1979) and Schön (1983) on Metacognition, Bereiter and Scardamalia's (1987) knowledge-transforming process, Facione's (1990) six Delphi skills, Risko and Gilbert (2016) on cognitive offloading, and Sundar's (2020) Theory of Interactive Media Effects.

The virtuous pathway: what fostered reasoning

GenAI supported Critical Thinking not by delivering knowledge but by creating conditions in which students had to reason about what it produced. The review names six enabling design features: evaluative engagement with AI-generated steps, error-detection tasks, multi-turn dialogue, prompt formulation before response, structured actionable feedback with a required response, and inquiry Simulation with embedded reasoning demands. In the strongest experiments the tool supplied material for judgment — Baroud et al. (2025) had 111 Moroccan baccalaureate biology students evaluate ChatGPT explanations of meiosis against textbook sources; Turós et al. (2025) applied Toulmin's argumentation model to essays from 452 Hungarian secondary students who argued against AI-generated claims and produced more backing, rebuttal and qualification. Kunnath and Botes (2025) embedded Magic School AI, LabXchange and Edpuzzle inside the 5E inquiry cycle (Engage, Explore, Explain, Elaborate, Evaluate) in South African science classrooms so that each phase demanded a different form of cognitive work, and Huang and Qiao (2024) found 136 Chinese secondary students in a STEAM-AI curriculum improved Computational Thinking given structured hypothesizing and iteration.

The mechanism is Scaffolding that raises rather than removes demand: GenAI absorbs mechanical effort (retrieval, first drafts, worked examples) so freed capacity can go to evaluation, comparison and argumentation — provided the tool fades as competence grows instead of performing the task permanently.

The vicious pathway: what eroded reasoning

Theme 2 studies show the mirror image. Uncritical acceptance arises when fluent, confident Large Language Models (LLMs) output signals reliability that the content does not warrant; the surface qualities that make a response readable are the cues that suppress scrutiny. Cognitive substitution occurs when AI supplies complete answers before a student has formulated the problem, skipping question formulation, hypothesis formation and evidence evaluation — Han and Han's (2025) AIStoryBot example, or the automatic personalization of PROSPETTIVA in Italian inclusive classrooms (Schicchi & Taibi, 2024) that reduced lower-performing students' engagement when they were not asked to express understanding first. Misaligned feedback breaks the revision cycle rather than activating it, and uncritical interaction is where individual habits harden: Kosmyna et al. (2025) reported reduced neural connectivity in executive-control regions after AI-assisted essay writing, and Castillo and Saavedra (2024) found habitual answer retrieval measurably impaired independent problem-solving in Peruvian secondary students.

The authors link this asymmetric risk to Prior Knowledge and developmental stage: students still building domain knowledge and epistemic judgment lack the Metacognition needed to notice that fluent output may be incomplete, biased or wrong, and Kuhn's (1999) developmental finding that reasoning dispositions formed in childhood persist makes early habituation especially consequential.

Stakeholders, the awareness gap, and the framework

Teachers, students and institutional stakeholders all converged on the same conditional view — GenAI supports reasoning when students must evaluate, question and push back, and undermines it when they need not — but they differed sharply in awareness. Teachers in Yin et al. (2025), Mwakalinga and Mabilika (2025), Tripathi et al. (2025), Santos et al. (2025, n = 379 Spanish secondary teachers) and Tongchai and Malakul (2025, n = 416 Thai teachers) named passive use, dependency and an implementation gap rooted in missing training rather than resistance; their concerns centered on Teacher AI Competency and the Educational Technology Developers who would need to build reason-preserving tools. Students, the group most at developmental risk, were least concerned, which the authors read as evidence that GenAI literacy cannot rely on self-awareness and must be explicit instruction in how outputs are generated and why they fail — including Hallucination Risk and Trust Calibration — embedded across subjects, not confined to a standalone module.

The Vicious or Virtuous framework organizes this over three interacting levels. Five categories of moderating factor — Human Factors, AI/System Factors, Contextual Conditions, Technological Affordances and Ethical and Societal Considerations — set the conditions; design choices then determine the pathway. The Virtuous Cycle runs on four principles: cognitive Scaffolding through guided constraints and question refinement; interaction design through multi-turn dialogue and user Learner Agency; feedback and adaptation that is specific, prioritized and actionable; and reflective and ethical design that asks how an output was produced and what might be missing. The Vicious Cycle inverts these into cognitive substitution, directive interaction, misaligned Feedback and uncritical interaction. Outcomes are temporally stratified — short-term cognitive-engagement patterns feed back into design, while long-term dispositions feed back into the moderating factors themselves, which is why the authors argue the K-12 window for intervention is real but closing.

Implications are assigned by role: teachers should design tasks requiring explanation, comparison, justification and revision of AI output and model critical interrogation openly; designers should build prompt-refinement gates, fading scaffolds, uncertainty displays and prioritized feedback; policymakers should make GenAI literacy a curriculum standard with assessed evaluation outcomes, require accountability for preserved student Learner Agency, and fund longitudinal work in under-resourced settings. The review's limits are its own evidence base: cross-sectional and short-horizon studies dominate, few track the same students over years, the neurocognitive findings are preliminary and not K–12-specific, and effect sizes come from a handful of quasi-experiments rather than a meta-analysis.

What this means for practice

  • Instructors. Require students to judge AI output before they can use it: the review's strongest reasoning gains came from tasks where learners evaluate, question and revise generated steps rather than copy them, as when 91 Grade 11 students in Jordan judged each step of a Wolfram Alpha solution across a six-week calculus unit.
  • Instructors. Build error-detection into the task on purpose, since asking 109 sixteen-year-old students to provoke and catch ChatGPT hallucinations led many of them to adopt "epistemic safeguarding" — restricting AI use to domains where they could verify answers.
  • Designers. Put the reasoning demand before the AI response: gate output behind prompt formulation and multi-turn dialogue, fade the Scaffolding as competence grows instead of letting it perform the task permanently, and prioritize Feedback so students must revise rather than merely read.
  • Designers. Surface uncertainty rather than fluency in the interface, because it was the confident, readable surface of Large Language Models (LLMs) output that suppressed scrutiny in the harm studies.
  • Researchers. Fund longitudinal, multi-site work in under-resourced settings: only 8 of the 38 studies documented harm, cross-sectional and short-horizon designs dominate, and the neurocognitive findings are preliminary and not K–12-specific.

Limitations

  • The review retained 38 studies from an initial pool of 1,392 records (35 from database screening plus 3 from a targeted hand search), and only 8 of those 38 documented harm, so the vicious pathway rests on a small subset of the evidence.
  • The search was restricted to English-language studies published between 2017 and 2025 and deliberately excluded higher-education, purely technical and non-critical-thinking work, which the authors note leaves the K–12 evidence base thinner than the higher-education literature and the theoretical grounding of the included studies inconsistent.
  • Short-horizon and cross-sectional designs dominate and few studies track the same students over years, so the claim that short-term engagement patterns harden into durable dispositions is an inference from developmental theory rather than an observed result.
  • The effect sizes come from a handful of quasi-experiments rather than a meta-analysis — 91 students in Jordan (η² = 0.21 to 0.35), 109 in Greece, 21 in Indonesia — and the supporting neurocognitive findings (Kosmyna et al., 2025) are preliminary and not specific to K–12 learners.

Connected Concepts

  • Critical Thinking — the outcome construct, defined through Facione's six Delphi skills (interpretation, analysis, evaluation, inference, explanation, self-regulation) and treated as disposition plus practice
  • Cognitive Offloading — the mechanism of harm when AI displaces rather than supports the cognitive effort reasoning requires
  • Generative AI — the technology whose effect the review argues is conditional on interaction design, not inherent
  • Large Language Models (LLMs) — the fluent, confident outputs whose surface qualities trigger uncritical acceptance in developing learners
  • AI Literacy — evaluative competence in how AI generates output and why it fails, shown to predict critical thinking and analytical skill
  • Scaffolding — Vygotskian support that must stay responsive and fade, contrasted with the "anti-scaffold" that performs the task entirely
  • Metacognition — self-monitoring that GenAI either externalizes for inspection or suppresses through fluent, authoritative-seeming output
  • Inquiry-Based Learning — the 5E cycle into which AI must be embedded at reasoning-demanding phases, not answer-supplying ones
  • Evaluative Judgment — the design requirement to judge AI output rather than copy it, the single most direct enabler of reasoning gains
  • AI Feedback Quality — structured, prioritized, actionable feedback versus dense unprioritised feedback that overwhelms revision
  • Prompt Engineering — prompt formulation before response, protected as the most cognitively demanding stage of inquiry
  • Human AI Collaboration — the distribution of cognitive responsibility between system and learner that determines the outcome

Connected Articles

Citation

Esmaeiligoujar, S., & Rahimi, S. (2026). Vicious or Virtuous? Designing Generative AI-Powered Learning Experiences to Foster Rather than Undermine Critical Thinking in K-12 Education. PsyArXiv preprint.

Embed this page

Copy the code below to embed a chromeless version of this page in a learning management system or other website. The embedded view hides the site header, navigation, and footer.