Li, Cui & Hagedorn (2026) PRISMA-review 67 empirical studies (2022β2025) on ChatGPT and university students' critical and creative thinking: effects are contingent on pedagogical framing, not the tool itself (generative-ai).
Li, Cui, and Hagedorn (2026) conducted a PRISMA-guided systematic review of 67 empirical studies (2022β2025) examining how ChatGPT influences university students' critical and creative thinking. Using a dual-lens framework β convergent (critical thinking) and divergent (creative thinking) processes β the review reveals that ChatGPT's cognitive effects are fundamentally contingent on pedagogical framing, not the tool itself.
Theoretical Framework
The review draws on an integrated constellation of theories spanning cognitive, sociocultural, and technological dimensions:
| Theory | Role in Framework |
|---|---|
| AI Literacy (Long & Magerko, 2020) | Foundational moderator β shapes prompt quality, epistemic vigilance, and verification practices |
| Self-Regulated Learning (Zimmerman, 2002) | Explains how structured engagement promotes planning, monitoring, and evaluative judgment |
| Cognitive Load Theory (Sweller, 2011) | Clarifies when ChatGPT reduces extraneous load (beneficial) vs. enables cognitive offloading (harmful) |
| Distributed Cognition (Hutchins, 1995) | Frames ChatGPT as an interactive cognitive artifact supporting iterative reasoning cycles |
| Cultural Historical Activity Theory (EngestrΓΆm, 1987) | Explains why outcomes differ across pedagogical ecosystems β rules, division of labor, task objects |
| Boundary Object Theory (Star & Griesemer, 1989) | Captures ChatGPT's interpretive flexibility across disciplines and tasks |
| Connectivism (Siemens, 2005) | Analyzes ChatGPT as a generative node in distributed learning networks |
This constellation treats ChatGPT as an interactive cognitive artifact within a socio-technical learning system β neither inherently beneficial nor detrimental, but contingent on how learners and instructors mobilize its affordances.
Study Characteristics
The 67 studies reflect the evolving nature of ChatGPT research:
- Design: 39% quantitative, 30% qualitative, 31% mixed-methods
- Disciplines: STEM 36%, Teacher Education 31%, Language/Writing 27%, Interdisciplinary 6%
- Geography: Asia 58%, Europe 18%, North America 10%, Oceania 4%, Africa 4%
- Publication trend: 10% in 2023, 58% in 2024, 31% in 2025 (through April)
- Focus: 34% CT only, 16% CrT only, 49% both
- Top venues: Education and Information Technologies (4), Frontiers in Education (3), JITE: Research (3), Computers & Education (3)
Methodological Note: Assessment Asymmetry
A critical methodological finding: CrT was more often assessed with direct performance tasks (TTCT, expert-rated artifacts), while CT relied more on indirect self-report measures. This asymmetry may partly explain why some studies report stronger evidence for CrT gains than CT gains, independent of actual cognitive effects.
Key Findings
Critical Thinking (56 studies)
Affordances (when ChatGPT was embedded in structured, scaffolded designs): 1. Metacognitive engagement (n=27) β Guided prompting, comparative analysis, and reflective writing enabled students to monitor thinking and exercise evaluative judgment 2. Argumentative structuring (n=22) β ChatGPT as dialogic scaffold or counterargument generator, especially with argument mapping and rubric-guided evaluation 3. Verification and error detection (n=19) β Fact-checking and triangulation behaviors emerged when students were explicitly taught to identify hallucinations 4. Self-regulated learning (n=17) β Structured prompts and revision cycles promoted planning, monitoring, and strategic adjustment 5. Disciplinary reasoning (n=15) β Complex analysis emerged organically when ChatGPT was a co-developer or critique target in authentic disciplinary tasks
Limitations (in unstructured contexts): 1. Cognitive offloading/overreliance (n=21) β Most common risk, especially among novice users and non-native speakers 2. Surface-level engagement (n=18) β Uncritical acceptance of AI outputs without appraisal 3. Erosion of argument development (n=14) β AI replaced the cognitive struggle integral to constructing ideas 4. Metacognitive offloading (n=12) β Surface-level edits without deeper planning or reflection 5. Epistemic boundary limits (n=10) β ChatGPT struggled with advanced rationality, logical consistency, and sustained Socratic dialogue
Creative Thinking (44 studies)
Affordances: 1. Ideation and divergent thinking (n=31) β Most frequently reported affordance; ChatGPT as brainstorming partner surfacing unique insights 2. Structural and expressive scaffolding (n=24) β Assisted with structuring content, experimenting with tone, stylistic expression 3. Dialogic engagement and perspective-shifting (n=18) β Functioned as co-designer in argument, debate, and role-based simulations 4. Affective and motivational activation (n=16) β Reduced creative anxiety; perceived as "brainstorming buddy" 5. Instructionally mediated gains (n=21) β Significant gains in originality, fluency, and elaboration when embedded in flipped classrooms or scaffolded creative modules (d=0.55β0.69 for key measures)
Limitations: 1. Creative passivity (n=20) β Diminished inclination to explore original ideas; substitution of cognitive effort 2. Loss of voice and affective authenticity (n=15) β AI outputs lacked individual style and emotional nuance 3. Suppression of iterative exploration (n=13) β Repetitive use, uncritical adoption, minimal conceptual recombination 4. Risk-related inhibition (n=11) β Privacy concerns and fear of underperformance suppressed creative risk-taking 5. Instructional deficit (n=14) β Without reflective prompts or structured interaction, ChatGPT functioned as passive answer provider
Co-Occurrence Patterns
The 33 studies examining both CT and CrT revealed three trajectories:
| Pattern | N | Description | Conditions |
|---|---|---|---|
| Synergistic Enhancement (CTβ, CrTβ) | 18 | Simultaneous gains in both domains | Inquiry-oriented tasks, scaffolded reflection, dialogic interaction, iterative refinement |
| Asymmetrical Development (CrTβ, CTβ) | 8 | Creative fluency improved but critical engagement declined | Unstructured use, emphasis on ideation over evaluation, minimal critical framing |
| Joint Cognitive Erosion (CTβ, CrTβ) | 4 | Both domains stagnated or declined | Passive/unscaffolded usage, task-completion focus, no metacognitive prompts |
The synergistic pattern aligns with critical-thinking-genai-scaffolding β when tasks require both generation and evaluation, AI supports both. The asymmetrical pattern is the most common failure mode: creativity flourishes at critical thinking's expense. The joint erosion pattern, while least common (n=4), is the most concerning β occurring when ChatGPT is used purely as a convenience tool with no pedagogical framing.
Discussion: ChatGPT as Cognitive Mediator
The review's central insight: ChatGPT functions as a cognitive mediator whose outcomes are contingent on how its affordances are mobilized. The most consistent divider was not ChatGPT's presence but the instructional ecology surrounding it:
- When tasks required verification, justification, and iterative revision, students treated AI outputs as provisional representations to interrogate
- When such norms were weak, fluent outputs lowered perceived task difficulty and encouraged premature closure
ChatGPT's semantic fluency emerged as both a strength and constraint β enabling rapid ideation while risking the masking of epistemic gaps. This maps onto metacognitive suppression risks: reduced perceived effort may encourage cognitive offloading rather than strategic load management.
The boundary object function β ChatGPT's interpretive flexibility across disciplines and tasks β supported cognitive adaptability and creative recombination, but also amplified surface-level synthesis when verification norms were absent. This connects to institutional norms and the broader activity systems in which AI tools operate.
Six Pedagogical Recommendations
1. Embed structured cognitive scaffolding β Stepwise activities: prompt design β output evaluation β iterative refinement. Frame ChatGPT as dialogic partner, not solution provider. This aligns with the six-process scaffolding framework and human-in-the-loop architectures.
2. Explicitly teach AI literacy for epistemic vigilance β Integrate modules on prompt refinement, hallucination recognition, bias detection, and contextual interpretation. Addresses the self-report vs. performance gap in AI evaluation skills.
3. Design tasks that co-activate CT and CrT through recursive inquiry β Open-ended case studies, argumentative writing with multi-perspective AI dialogue, project-based tasks requiring both generation and analytical reflection. This is the practical implementation of the dual-lens framework.
4. Implement reflection protocols for cognitive regulation β Guided prompts after each interaction: "What was most useful/misleading?", "How did this shape your thinking?", "What would you change in your next prompt?" Reinforces metacognitive monitoring.
5. Leverage ChatGPT as a connective node for interdisciplinary thinking β Cross-domain tasks that draw on ChatGPT's broad knowledge while critically examining disciplinary assumptions. Supports dialogic partner and connectivist learning.
6. Position feedback as a multi-source process β Triangulate AI feedback with peer review, instructor input, and self-assessment. Creates multi-source feedback loops that mitigate overreliance.
Limitations of the Review
- English-language, peer-reviewed journal articles only β excludes conference proceedings (LAK, AIED, L@S) and non-English research
- Dominance of Asian institutions (58%) and early-adopter settings limits generalizability
- Most studies were cross-sectional/short-term; no longitudinal tracking of cognitive habit formation
- Assessment asymmetry: CrT measured with performance tasks, CT with self-reports β apparent robustness differences may reflect measurement, not reality
- Rapidly evolving technology β findings tied to specific ChatGPT versions; living systematic reviews needed
- Publication bias likely favors positive findings in this emerging field
Implications for the Wiki
This review is a keystone synthesis connecting multiple threads in the AI education evidence base:
- critical-thinking-genai-scaffolding shares the core premise β pedagogical design determines whether GenAI helps or harms thinking β and the six-process framework maps directly onto this review's scaffolding recommendations
- metacognition identifies the same suppression risk and regulatory mechanisms the review documents at scale across 67 studies
- ai-literacy-assessment-misalignment explains why students struggle to calibrate AI evaluation β the review confirms this as a critical moderator
- human-in-the-loop-ai provides implementation architectures for the scaffolding strategies recommended
- ai-learning-companions-framework offers design paradigms aligned with "dialogic partner" and "boundary object" concepts
- feedback-loop operationalizes the multi-source feedback recommendation
- faculty-development-genai is essential β educators need training to implement these scaffolds
- student-experience captures the learner perspective on usage patterns
- higher-ed and universities-ai-era-rethinking provide the institutional context
- institutional-change-framework-ai frames how activity systems must adapt
The review's core insight β that ChatGPT's cognitive effects are contingent on pedagogy, not inherent to the technology β reinforces a pattern visible across the wiki: AI in education succeeds or fails based on how it is implemented, not what it can do.
Related Pages
- critical-genai-use-predictors β Disposition toward critical thinking predicts critical use
- critical-thinking-genai-scaffolding β Scaffolding framework for critical thinking with GenAI; directly complementary
- metacognition β Metacognitive regulation and suppression risks confirmed at scale
- higher-ed β Higher education context
- ai-literacy-assessment-misalignment β Self-report vs. performance gap in AI evaluation
- human-in-the-loop-ai β Human oversight architectures for educational AI
- ai-learning-companions-framework β Designing AI as dialogic partner
- faculty-development-genai β Educator training for AI integration
- feedback-loop β Multi-source feedback in AI learning environments
- student-experience β Learner perspectives on AI tools
- universities-ai-era-rethinking β Institutional transformation for AI era
- institutional-change-framework-ai β Framework for AI adoption in institutions