On this page

Critical thinking — the ability to analyze, evaluate, and synthesize information — is both a skill that AI tools can help develop and a competency that students must apply when using AI. In AI in education research, critical thinking appears in two interrelated forms: as a learning objective (teaching students to think critically) and as a safeguard against uncritical AI reliance.

Questions to Consider

  • How confident are you in your ability to spot a false or misleading AI-generated answer? Research suggests self-reported AI competence far exceeds actual evaluation ability — how would you test yourself?
  • Critical thinking here appears in two forms: a skill to teach, and a safeguard against uncritical reliance on AI. Can you think of a situation where a tool that 'teaches' critical thinking is actually training its opposite?
  • One study found that having students interrogate AI-generated mistakes produced large gains in higher-order thinking. How might deliberately exposing errors — rather than hiding them — be a more powerful teaching move than you assumed?
  • Easy access to AI answers can displace critical engagement before students realize it. What design feature, rather than a policy or a ban, could keep the cognitive effort alive?
  • AI advice has been shown to suppress the willingness to say 'I don't know' — even when the advice is wrong. How does that change what it means to create a classroom culture where questioning is safe?

Introduction

Critical thinking is central to AI Literacy — students who cannot critically evaluate AI outputs are vulnerable to Over-Reliance, hallucinated information, and biased recommendations. Research on Cognitive Offloading shows that easy access to AI answers can displace critical engagement, while Socratic approaches that withhold direct answers preserve the cognitive effort necessary for deeper thinking.

Critical thinking in AI education research

The knowledge base's articles explore critical thinking through design-based and empirical lenses. Adversarial AI agents enact constructive conflict to prompt reconsideration in novice designers — a Socratic variant that forces critical re-evaluation. RCT research on GenAI in teaching raises the question of whether AI tools that optimize for surface-level outcomes may inadvertently suppress the critical thinking that leads to deeper learning. Measured evidence sharpens the point: whether critical thinking moves depends on how AI-mediated feedback and tasks are designed, and on learners' metacognitive regulation, not on access to a model.

Reviews of ChatGPT's impact on thinking document mixed findings: AI can scaffold critical analysis when used deliberately (e.g., asking students to critique AI-generated arguments), but it can also short-circuit thinking when used as an answer engine. This tension connects to How to Assess AI Literacy: Misalignment Between Self-Reported and Objective-Based Measures research showing that self-reported AI competence far exceeds actual critical evaluation ability.

  • Higher-order cognitive engagement in student-AI chat. Chang and Li (2026) find that ~62% of student prompts to AI encode higher-order cognitive demand, with Bloom-level profiles varying by discipline (STEM Apply-prevalent 20.8%, language Understand-prevalent 31.7%, social science Create-prevalent 33.8%). Their within-person design shows the same students produce significantly more higher-order prompts in social science than STEM courses (p < .001), indicating that disciplinary context shapes critical and higher-order engagement with AI.

  • AI as a catalyst for critical media literacy in children. Demir and Akar (2026) evaluate an 18-hour, 5E-model critical media literacy program for fourth-grade Turkish students in which generative AI (ChatGPT, Grammarly) acted as a pedagogical agent embedded phase-by-phase rather than an add-on. Paired-samples comparisons showed large gains in media reading (+3.50), writing (+1.67), and total media literacy (+5.17, all p < .01), with between-group post-test effect sizes of Cohen's d = 1.12 (reading), 1.18 (writing), and 1.31 (total literacy) favoring the AI-supported group. Qualitative analysis (interviews, student posters/drawings/slogans, classroom observation) surfaced six domains of critical media literacy growth — digital self-protection and data privacy, purposeful and responsible media use, safe communication and boundary awareness, critical evaluation and misinformation awareness, online risk awareness, and media Ethics/digital citizenship — indicating that deliberately interrogating AI-mediated content can cultivate critical analysis and reflection in young learners.

  • Dimension-specific critical-thinking gains in primary multimodal writing. Lu et al. (2027) followed 60 Grade 5 students through an eight-week Conversational AI-supported multimodal writing practice in which they turned narratives into AI-generated images and short videos. Repeated-measures analysis across six critical-thinking dimensions found sustained gains (T1→T2 and T1→T3) in interpretation, analysis, evaluation, and explanation, a short-lived self-AI Regulation in Education gain, and no change in inference — an uneven, dimension-level pattern that an aggregate critical-thinking score would have hidden. The authors argue the AI-generated visuals externalized meaning and thereby lowered the inferential demand writing normally imposes, while peer collaboration (peer questions that forced inferring others' interpretations) supplied the occasions for inference the solo AI interaction did not. The design lesson: multimodal AI composing supports several critical-thinking facets but should be paired with continued Scaffolding and structured peer exchange to preserve inference and self-regulation.

  • AI scaffolding and offloading pull critical thinking in opposite directions. Davor, Larbi and Boateng (2026) surveyed 533 university students in Ghana and found that AI task scaffolding predicted higher critical thinking (β = .185) while cognitive offloading tendency predicted lower critical thinking (-.240); AI verification literacy had no direct effect on critical thinking and worked only through metacognitive self-regulation, a full mediation pattern the authors read as evidence that teaching students to fact-check AI is not enough on its own. (Davor et al. 2026)

  • Dependence, not use, is where the association with critical thinking turns. Shojaei and colleagues (2026) surveyed 412 business students in Oman and found a near-zero bivariate correlation between GenAI use and self-reported critical-thinking disposition (r = 0.050), with dependence predicting lower disposition (β = -0.389) and weakening the link from use to disposition (β = -0.239), so that the simple slope fell from 0.424 at low dependence to -0.054 at high dependence. (Shojaei et al. 2026)

  • A short reflection prompt makes reliance on AI advice more discriminative. In a three-condition experiment with 342 undergraduates, Ren (2026) found that open ChatGPT support produced acceptance of incorrect AI recommendations on 62.4% of trials, falling to 39.7% with a brief metacognitive reflection prompt (OR = 0.40, 95% CI [0.28, 0.56]); reflection also improved awareness calibration (0.59 vs. 0.41) and cut the AI-specific attribution bias index from 0.42 to 0.21 without reducing recommendation accuracy or triggering blanket rejection of useful advice. (Ren 2026)

  • Teacher-curated ChatGPT feedback raises critical thinking through higher-order revision. Chen and colleagues (2026) ran an 18-week quasi-experiment with 64 undergraduates, all pre-service chemistry, physics or mathematics teachers, comparing conventional teacher feedback (n = 32) against ChatGPT-assisted teacher feedback (n = 32) across two argumentative writing tasks. The groups began equivalent, and only the assisted group improved significantly (p < 0.001), with Cohen's d rising from 0.25 at pretest to 3.74 at posttest. The mechanism was the kind of feedback, not the mere presence of a model: the assisted group received more demonstration (33.1% vs. 10.9%) and closely questioning (21.5% vs. 6.0%) feedback, revised 91.00% (435 of 478) of feedback units against 85.21% (242 of 284), and epistemic network analysis linked those revisions to analysis, evaluation and creation rather than recognition and understanding, while teacher-only instruction and evaluation feedback mostly produced recognition and understanding revisions. Teachers treated the model's output as draft material, expanding, revising or discarding it (14% discarded, 9% needing correction), and six of eight interviewees still flagged imprecision or unprofessionalism. The design lesson is that the critical-thinking gain came from feedback that models alternatives and interrogates the student's reasoning, with a teacher curating what the model produced; the small, single-semester sample means the very large effect size should be read with caution. (Chen et al. 2026)

  • A semester of AI access left critical thinking flat, while reflective use tracked it. Melanou, Beege and Kimmig (2026) followed 87 business informatics students through a nine-week course in three parallel conditions (tutor-framed AI, unguided AI, and no AI) measured three times. Knowledge rose in every group with no advantage for either AI condition and no Matthew effect (BF01 = 8.70), while self-reported critical thinking and Motivation stayed stable. What separated students was reflective use, the practice of checking sources and verifying AI output before adopting it: it was higher in the AI condition than in the control group (means of 3.72 vs. 2.82) and predicted critical thinking at the final measurement (R² = 0.183, β = 0.43, p < 0.001). The result qualifies any expectation that a course of AI use will move critical thinking on its own: metacognitive regulation, not tool access, is where the association lives, and the authors note a semester is likely too short to see durable change. (Melanou et al. 2026)

Connections to other concepts

Critical thinking intersects with Scaffolding (designing AI support that maintains cognitive demand), Prompt Engineering (formulating questions that elicit critical analysis), and Over-Reliance (knowing when to trust and when to question AI). It is foundational to Academic Integrity and serves as a key dimension of AI Literacy frameworks across both K-12 and Higher Education contexts.

  • AI errors as provocations for higher-order thinking: Hosseini (2026) operationalizes Bloom's higher-order levels (Analyze, Evaluate, Create) by having students interrogate AI-generated mistakes in a database course, with significant pre/post gains (Cohen's d=1.49) in subject-matter competency.

  • Two-sided auditing of AI explanations. Bernstein and Sibia (2026) used Paul–Elder standards (accuracy, clarity, assumptions, point of view) as interview probes with ten students who had completed CS2, and found mechanism-level scrutiny of GenAI explanations: students located where an analogy's mapping broke (an island-route analogy for a linked list that implied a circle, a badminton rally offered for recursion that had no guaranteed shrinking input), demanded precise wording over hedging, and treated explanations as arguments carrying a point of view. Crucially, that scrutiny tracked source- or target-domain expertise rather than personal interest — reframing critical evaluation of AI output as a knowledge problem ("two-sided analogy auditing") rather than a dispositional one — and suggesting that assigning flawed AI analogies as objects to inspect and repair is a harder check on conceptual understanding than reading a finished explanation.(Flawed but Memorable: Student Critical Reception of Interest-Personalized GenAI Analogies in Computing Education)

  • Critical thinking as non-outsourceable engagement. Xie (2026) adds a philosophical counterpart from Daoist self-cultivation: because AI is an opaque "black box," a framework oriented to harmonizing uncertainty rather than adjudicating truth in absolute terms better suits the current epistemic landscape, and critical thinking becomes sustained, first-person, non-outsourceable engagement with reality rather than a demonstrable rational procedure. In Neidan (內丹) practice "there are no cognitive shortcuts," so AI is positioned as "not a cognitive surrogate but an instrumental adjunct" to human flourishing.(Alternative AI Philosophy: Daoism as Method for AI in Education)

  • Role rotation as a structure for critical human-AI interaction. Kenzhebayeva and colleagues (2026) report a design-based study in which 62 pre-service educational psychologists rotated through four professional roles (Case Constructor, Research Analyst, Practitioner-Interventionist, Reflective Researcher) that made generated recommendations the object of discussion: participants compared AI output with psychological theory and modified or rejected recommendations that did not fit the case, and later cycles showed more requests for theoretical justification, while overreliance on apparently authoritative responses persisted. (Kenzhebayeva et al. 2026)

  • Verification-centered integration in a discipline. A critical review of generative AI in university chemistry education (Vega-Baudrit and Rivera Álvarez, 2026) argues that because chemical reasoning must be coordinated across macroscopic, submicroscopic, and symbolic representations, students cannot verify what they do not understand, so Prior Knowledge and Scaffolding come first and verification should be designed into Assessment as an assessed activity: identify a false assumption, correct a unit or mechanism error, or justify rejecting a generated answer, keeping prompt logs and revision histories as reasoning traces. (Vega-Baudrit and Rivera Álvarez 2026)

Connected Concepts

Connected Articles

Connected FAQs

Embed this page

Copy the code below to embed a chromeless version of this page in a learning management system or other website. The embedded view hides the site header, navigation, and footer.