Research Article
Can Large Language Models Foster Critical Thinking, Teamwork, and Problem-Solving Skills in Higher Education?: A Literature Review
Synthesis: Can Large Language Models Foster Critical Thinking, Teamwork, and Problem Solving Skills in Higher Education?: A Literature Review
Key Findings
- Systematic literature review following the PRISMA 2020 protocol, searching the Web of Science Core Collection for studies published 2023–2024; of 203 studies screened, 22 articles were included (10 from search string A, 12 from search string B), and the review is registered in PROSPERO (CRD420251165731).
- LLMs often produce incomplete or incorrect responses, prompting students to question, verify, and improve the information — creating validation-and-correction cycles that the reviewed studies link to enhanced critical thinking and mental independence.
- LLMs act as catalysts for collaboration: supporting idea generation, organization, and peer feedback, simulating rubric-based evaluations and expert reviews, and democratizing access to knowledge — promoting more equitable Collaborative Learning, especially in diverse and large-class settings.
- For problem-solving, LLMs help students explore alternative solutions, incorporate interdisciplinary perspectives, and simulate authentic real-world scenarios — e.g., clinical and ethical dilemmas in health sciences, and prototype testing, debugging, and refinement in STEM.
- LLMs deliver scalable, timely, personalized feedback in large courses (near-instant responses reduce delay and dependence on instructor availability) and can reduce faculty workload by automating feedback on routine tasks such as drafts, quizzes, and structured assignments.
- LLMs support assessment reform: generating rubrics and varied assessment items aligned with learning outcomes, and shifting assessment from factual recall toward scenario-based, competency-based evaluation.
About the Review
The review addressed four research questions (RQ1–RQ4) on how LLMs can foster critical thinking, collaborative skills, and problem-solving; bridge theory and practice; deliver personalized feedback in large courses; and support assessment development. Search criteria limited records to peer-reviewed English-language journal articles published between January 2023 and December 2024 (IC1–IC4), with exclusions for duplicates, non-journal document types, non-English language, and early access (EC1–EC4). Of 165 records entering the inclusion/exclusion stage, 117 were excluded and 48 papers were retrieved for eligibility assessment against five criteria (C1–C5): explicit teaching–learning focus, direct use of LLMs by students, pedagogical or subject-specific application in higher education, practical or empirical evidence of learning outcomes, and structured data or evaluation. The 48 eligible documents spanned 35 countries — led by the United States (n = 9), Germany (n = 8), Australia (n = 6), and England (n = 6) — with Frontiers Media SA and MDPI each publishing seven of the analyzed documents. Methodological rigor was enforced through the PRISMA 2020 protocol, the structured eligibility criteria, and PROSPERO registration.
Main Findings
- Critical thinking: students learn to evaluate the accuracy, consistency, and trustworthiness of Large Language Models (LLMs)-generated content; effectively framing prompts improves their understanding and analytical skills, engaging them in comparison, synthesis, and assessment.
- Collaboration: LLMs enrich teamwork by aiding idea generation and peer feedback, simulating expert or peer roles from various fields, and lowering participation barriers in large and diverse classes.
- Theory–practice gap: simulations, hands-on project work, and instant alternative solutions on error accelerate learning cycles and support self-management and independent learning via anytime/anywhere access.
- Feedback at scale: LLM-supported feedback extends beyond one-way correction to become a social, iterative process, and enables students to generate practice tests and self-assessments that strengthen metacognitive skills.
- Assessment: scenario-based and problem-focused evaluation lets students demonstrate applied skills in realistic settings (engineering, health care, business), shifting the assessment culture from knowledge reproduction to applied understanding.
- Systemic implications: the authors argue LLMs can transform curricula, feedback, and assessment at institutional and policy levels — improving educational equity, workforce readiness, and innovation — provided institutions establish ethical, equitable, and sustainable adoption frameworks.
What this means for practice
- Instructors. Turn model error into coursework: assign students to question, verify, and correct LLM output, the validation-and-correction cycle the reviewed studies link to Critical Thinking and mental independence.
- Instructors. Use LLMs to generate practice tests, self-assessments, and rubric-based peer feedback so feedback becomes an iterative, social process rather than one-way correction — a move the review ties to stronger metacognitive skill.
- Administrators. Deploy LLMs where feedback delay is the binding constraint, in large-enrollment courses, and pair that with GenAI literacy in curricula and explicit faculty preparation for responsible use.
- Policymakers. Treat the reported benefits as conditional rather than automatic: the authors argue the impact depends on institutions establishing ethical, equitable, and sustainable adoption frameworks.
- Researchers. Extend the evidence base beyond the review's window — its 22 included articles were published in 2023–2024 only, and no included document examined adoption challenges such as ethics or privacy.
Limitations
The authors explicitly list four limitations: (a) the search relied on the Web of Science Core Collection without considering other databases; (b) inclusion was restricted to scientific journal publications, excluding other document types; (c) the included documents did not focus on challenges of LLM adoption in teaching and learning, such as ethical concerns or privacy; and (d) the review covered only 2023 and 2024, the first two years of the technology's emergence. No formal risk-of-bias or quality-assessment instrument was applied to the included studies beyond the PRISMA 2020 protocol and the structured eligibility criteria.
Citation
Martínez-Peláez, R., Mena, L. J., Toral-Cruz, H., Ochoa-Brust, A., González Potes, A., Flores, V., Ostos, R., Ramírez Pacheco, J. C., Félix, R. A., & Félix, V. G. (2025). Can large language models foster critical thinking, teamwork, and problem-solving skills in higher education? A literature review.