Research Article
The cognitive impact of ChatGPT in higher education: A systematic review of critical and creative thinking outcomes
Synthesis: Li, Cui, and Hagedorn (2026) conducted a PRISMA-guided systematic review of 67 empirical studies (2022–2025) examining how ChatGPT influences university students' critical and creative thinking. Using a dual-lens framework — convergent (critical thinking) and divergent (creative thinking) processes — the review reveals that ChatGPT's cognitive effects are fundamentally contingent on pedagogical framing, not the tool itself.
Theoretical Framework
The review draws on an integrated constellation of theories spanning cognitive, sociocultural, and technological dimensions:
| Theory | Role in Framework |
|---|---|
| AI Literacy (Long & Magerko, 2020) | Foundational moderator — shapes prompt quality, epistemic vigilance, and verification practices |
| Self-Regulated Learning (Zimmerman, 2002) | Explains how structured engagement promotes planning, monitoring, and evaluative judgment |
| Cognitive Load Theory (Sweller, 2011) | Clarifies when ChatGPT reduces extraneous load (beneficial) vs. enables cognitive offloading (harmful) |
| Distributed Cognition (Hutchins, 1995) | Frames ChatGPT as an interactive cognitive artifact supporting iterative reasoning cycles |
| Cultural Historical Activity Theory (Engeström, 1987) | Explains why outcomes differ across pedagogical ecosystems — rules, division of labor, task objects |
| Boundary Object Theory (Star & Griesemer, 1989) | Captures ChatGPT's interpretive flexibility across disciplines and tasks |
| Connectivism (Siemens, 2005) | Analyzes ChatGPT as a generative node in distributed learning networks |
This constellation treats ChatGPT as an interactive cognitive artifact within a socio-technical learning system — neither inherently beneficial nor detrimental, but contingent on how learners and instructors mobilize its affordances.
Study Characteristics
The 67 studies reflect the evolving nature of ChatGPT research:
- Design: 39% quantitative, 30% qualitative, 31% mixed-methods
- Disciplines: STEM 36%, Teacher Education 31%, Language/Writing 27%, Interdisciplinary 6%
- Geography: Asia 58%, Europe 18%, North America 10%, Oceania 4%, Africa 4%
- Publication trend: 10% in 2023, 58% in 2024, 31% in 2025 (through April)
- Focus: 34% CT only, 16% CrT only, 49% both
- Top venues: Education and Information Technologies (4), Frontiers in Education (3), JITE: Research (3), Computers & Education (3)
Methodological Note: Assessment Asymmetry
A critical methodological finding: CrT was more often assessed with direct performance tasks (TTCT, expert-rated artifacts), while CT relied more on indirect self-report measures. This asymmetry may partly explain why some studies report stronger evidence for CrT gains than CT gains, independent of actual cognitive effects.
Key Findings
Critical Thinking (56 studies)
Affordances (when ChatGPT was embedded in structured, scaffolded designs):
- Metacognitive engagement (n=27) — Guided prompting, comparative analysis, and reflective writing enabled students to monitor thinking and exercise evaluative judgment
- Argumentative structuring (n=22) — ChatGPT as dialogic scaffold or counterargument generator, especially with argument mapping and rubric-guided evaluation
- Verification and error detection (n=19) — Fact-checking and triangulation behaviors emerged when students were explicitly taught to identify hallucinations
- Self-regulated learning (n=17) — Structured prompts and revision cycles promoted planning, monitoring, and strategic adjustment
- Disciplinary reasoning (n=15) — Complex analysis emerged organically when ChatGPT was a co-developer or critique target in authentic disciplinary tasks
Limitations (in unstructured contexts):
- Cognitive offloading/overreliance (n=21) — Most common risk, especially among novice users and non-native speakers
- Surface-level engagement (n=18) — Uncritical acceptance of AI outputs without appraisal
- Erosion of argument development (n=14) — AI replaced the cognitive struggle integral to constructing ideas
- Metacognitive offloading (n=12) — Surface-level edits without deeper planning or reflection
- Epistemic boundary limits (n=10) — ChatGPT struggled with advanced rationality, logical consistency, and sustained Socratic dialogue
Creative Thinking (44 studies)
Affordances:
- Ideation and divergent thinking (n=31) — Most frequently reported affordance; ChatGPT as brainstorming partner surfacing unique insights
- Structural and expressive Scaffolding (n=24) — Assisted with structuring content, experimenting with tone, stylistic expression
- Dialogic engagement and perspective-shifting (n=18) — Functioned as co-designer in argument, debate, and role-based simulations
- Affective and motivational activation (n=16) — Reduced creative anxiety; perceived as "brainstorming buddy"
- Instructionally mediated gains (n=21) — Significant gains in originality, fluency, and elaboration when embedded in flipped classrooms or scaffolded creative modules
Limitations:
- Creative passivity (n=20) — Diminished inclination to explore original ideas; substitution of cognitive effort
- Loss of voice and affective authenticity (n=15) — AI outputs lacked individual style and emotional nuance
- Suppression of iterative exploration (n=13) — Repetitive use, uncritical adoption, minimal conceptual recombination
- Risk-related inhibition (n=11) — Privacy concerns and fear of underperformance suppressed creative risk-taking
- Instructional deficit (n=14) — Without reflective prompts or structured interaction, ChatGPT functioned as passive answer provider
Co-Occurrence Patterns
The 33 studies examining both CT and CrT revealed three trajectories:
| Pattern | N | Description | Conditions |
|---|---|---|---|
| Synergistic Enhancement (CT↑, CrT↑) | 18 | Simultaneous gains in both domains | Inquiry-oriented tasks, scaffolded reflection, dialogic interaction, iterative refinement |
| Asymmetrical Development (CrT↑, CT↓) | 8 | Creative fluency improved but critical engagement declined | Unstructured use, emphasis on ideation over evaluation, minimal critical framing |
| Joint Cognitive Erosion (CT↓, CrT↓) | 4 | Both domains stagnated or declined | Passive/unscaffolded usage, task-completion focus, no metacognitive prompts |
The synergistic pattern aligns with Scaffolding critical thinking with generative AI: Design principles for integrating large language models in higher — when tasks require both generation and evaluation, AI supports both. The asymmetrical pattern is the most common failure mode: creativity flourishes at critical thinking's expense. The joint erosion pattern, while least common (n=4), is the most concerning — occurring when ChatGPT is used purely as a convenience tool with no pedagogical framing.
Discussion: ChatGPT as Cognitive Mediator
The review's central insight: ChatGPT functions as a cognitive mediator whose outcomes are contingent on how its affordances are mobilized. The most consistent divider was not ChatGPT's presence but the instructional ecology surrounding it:
- When tasks required verification, justification, and iterative revision, students treated AI outputs as provisional representations to interrogate
- When such norms were weak, fluent outputs lowered perceived task difficulty and encouraged premature closure
ChatGPT's semantic fluency emerged as both a strength and constraint — enabling rapid ideation while risking the masking of epistemic gaps. This maps onto metacognitive suppression risks: reduced perceived effort may encourage cognitive offloading rather than strategic load management.
The boundary object function — ChatGPT's interpretive flexibility across disciplines and tasks — supported cognitive adaptability and creative recombination, but also amplified surface-level synthesis when verification norms were absent. This connects to institutional norms and the broader activity systems in which AI tools operate.
Six Pedagogical Recommendations
- Embed structured cognitive scaffolding — Stepwise activities: prompt design → output evaluation → iterative refinement. Frame ChatGPT as dialogic partner, not solution provider. This aligns with the six-process scaffolding framework and human-in-the-loop architectures.
- Explicitly teach AI literacy for epistemic vigilance — Integrate modules on prompt refinement, hallucination recognition, bias detection, and contextual interpretation. Addresses the self-report vs. performance gap in AI evaluation skills.
- Design tasks that co-activate CT and CrT through recursive inquiry — Open-ended case studies, argumentative writing with multi-perspective AI dialogue, project-based tasks requiring both generation and analytical reflection. This is the practical implementation of the dual-lens framework.
- Implement reflection protocols for cognitive AI Regulation in Education — Guided prompts after each interaction: "What was most useful/misleading?", "How did this shape your thinking?", "What would you change in your next prompt?" Reinforces metacognitive monitoring.
- Leverage ChatGPT as a connective node for interdisciplinary thinking — Cross-domain tasks that draw on ChatGPT's broad knowledge while critically examining disciplinary assumptions. Supports dialogic partner and connectivist learning.
- Position feedback as a multi-source process — Triangulate AI feedback with peer assessment, instructor input, and self-assessment. Creates multi-source feedback loops that mitigate overreliance.
Implications for the Knowledge Base
This review is a keystone synthesis connecting multiple threads in the AI education evidence base:
- Scaffolding critical thinking with generative AI: Design principles for integrating large language models in higher shares the core premise — pedagogical design determines whether GenAI helps or harms thinking — and the six-process framework maps directly onto this review's scaffolding recommendations
- Metacognition identifies the same suppression risk and regulatory mechanisms the review documents at scale across 67 studies
- How to Assess AI Literacy: Misalignment Between Self-Reported and Objective-Based Measures explains why students struggle to calibrate AI evaluation — the review confirms this as a critical moderator
- Human-in-the-Loop provides implementation architectures for the scaffolding strategies recommended
- Building AI Companions that Prioritise Learning over Performance offers design paradigms aligned with "dialogic partner" and "boundary object" concepts
- Feedback Loop operationalizes the multi-source feedback recommendation
- Educational Development is essential — educators need training to implement these scaffolds
- Student Experience captures the learner perspective on usage patterns
- Higher Education and The University AI Didn''t Replace: Rethinking Universities in the AI Era provide the institutional context
- A Framework for Institutional Change in the Age of AI frames how activity systems must adapt
The review's core insight — that ChatGPT's cognitive effects are contingent on pedagogy, not inherent to the technology — reinforces a pattern visible across the knowledge base: AI in education succeeds or fails based on how it is implemented, not what it can do.
What this means for practice
- Instructors. Require verification, justification, and iterative revision on every ChatGPT task. When those norms were weak, fluent output lowered students' perceived task difficulty and encouraged premature closure.
- Instructors. Build assignments that force generation and evaluation into the same task — open-ended case analysis with multi-perspective AI dialogue, for instance — so students land in the synergistic trajectory rather than the asymmetrical one, where creative fluency improves while critical engagement declines.
- Instructors. Front-load AI literacy instruction on prompt refinement, hallucination recognition, and bias detection. The review treats AI literacy as the moderator shaping prompt quality and verification behavior, not as optional orientation content.
- Researchers. Shift the comparison from ChatGPT's presence versus absence to the instructional ecology around it — task structure, verification norms, and feedback sources are the variables the reviewed evidence ties to divergent outcomes.
Limitations
- The synthesis rests on 67 English-language, peer-reviewed journal articles only, excluding conference proceedings (LAK, AIED, L@S) and non-English research; the review notes this restriction likely favors positive findings in an emerging field.
- Geographical concentration: 58% of the studies came from Asia and 10% from North America, so the evidence base leans toward early-adopter settings and generalizes unevenly across higher education systems.
- Assessment asymmetry: creative thinking was more often evaluated with direct performance tasks while critical thinking relied on indirect self-report measures, so apparent differences in robustness across the two domains partly reflect how each was operationalized and measured.
- Nearly all included studies were cross-sectional or short-term, with no longitudinal tracking of cognitive habit formation, and the technology changes faster than the literature — the review calls for living systematic reviews tied to specific ChatGPT versions.
Citation
Li, C., Cui, H., & Hagedorn, L. S. (2026). The cognitive impact of ChatGPT in higher education: A systematic review of critical and creative thinking outcomes. Computers and Education: Artificial Intelligence.