Research Article
Generative AI in Higher Education: A Systematic Review of Opportunities, Challenges, and Pedagogical Innovations (2022–2025)
Synthesis: This PRISMA 2020-guided systematic review synthesizes 125 peer-reviewed studies (2022–2025) on Generative AI in Higher Education, documenting exponential adoption (92% student usage by 2025, up from 66% in 2024), four primary application domains, and persistent challenges around Academic Integrity, algorithmic Bias Mitigation, Hallucination Risk, Educational Development gaps, and Equity/Digital Divide concerns. The authors propose an integrated four-dimensional framework — pedagogical integration, AI Literacy development, ethical AI Governance, and systemic support — for responsible GenAI implementation, emphasizing that context-sensitive, holistic adoption matters more than technological provision alone.
Key Findings
- 125 studies synthesized via PRISMA 2020. Empirical evidence was screened from 2,847 records across Scopus, Web of Science, ERIC, IEEE, and Google Scholar (with PubMed for health-sciences education), yielding 125 included studies; 18.4% originated from Global South contexts across 42 countries.
- Exponential GenAI adoption. Student usage reached 92% by 2025 — a 26-point jump from 66% in 2024 — with publication volume itself growing from 9 studies in 2022 to 51 in 2024.
- Four application domains. Automated Feedback and assessment (53.6% of studies), personalized learning support and tutoring (44.8%), critical skill development (30.4%), and research/administrative support (24.8%).
- Persistent challenges. Academic Integrity concerns dominate (71.2%), compounded by flawed detection tools (30–50% false-positive rates), 10–40% hallucination rates, 60–70% of faculty feeling unprepared, and equity gaps driven by differential access to premium tools.
- Emergent pedagogical innovations. AI Literacy frameworks, AI-adapted Technological Pedagogical Content Knowledge (TPACK) models, assessment redesign toward process documentation, participatory policy development, and Prompt Engineering curricula offer pathways for responsible integration.
Background and Method
Following the public release of ChatGPT by OpenAI in November 2022, advanced large language models democratized generative AI access for students, educators, and researchers worldwide. The review argues that GenAI's transformative potential — real-time personalized tutoring, automated formative feedback, multilingual translation, and code generation — is coupled with equally novel risks, including content hallucination, algorithmic bias reflecting training-data inequities, and fundamental questions about authorship and original thought in academic contexts.
Methodologically, the review followed PRISMA 2020 reporting guidelines, searched six databases (Scopus, Web of Science, ERIC, IEEE Xplore, PubMed, Google Scholar) between May 15–30 2026 for publications spanning January 2022 through December 2025, and screened 2,224 unique records after deduplication. Full-text review of 348 records excluded 223 articles, leaving 125 studies. Data extraction used a standardized pilot-tested form, quality was appraised with the Mixed Methods Appraisal Tool and AMSTAR 2, and synthesis employed narrative thematic analysis coded in NVivo 14. Studies concentrated in STEM and language education contexts, predominantly undergraduate samples (54.4%), with survey research (37.6%) the most common methodology.
Applications of GenAI in Higher Education
The thematic analysis surfaced four primary application domains. Automated feedback and assessment (67 studies) covered personalized draft feedback, generation of practice problems and quiz items, rubric-based evaluation of short answers, and real-time Formative Assessment via chatbots. Evidence on effectiveness was mixed: GenAI feedback matched instructor feedback in technical accuracy but lacked affective elements such as encouragement and recognition of individual growth trajectories. Personalized learning support and tutoring (56 studies) included on-demand concept tutoring, scaffolded research-paper development, language-learning assistance for ESL/EFL students, and Accessibility accommodations for students with disabilities. Longitudinal cases reported gains in Self-Efficacy, engagement, persistence, and Help-Seeking, alongside emerging concerns about over-reliance that inhibited metacognitive development and peer collaboration.
Critical skill development (38 studies) positioned GenAI as a catalyst for higher-order thinking rather than knowledge delivery, through critical evaluation of AI output, creative divergent-thinking exercises, Prompt Engineering assignments, and collaborative human-AI workflows. These approaches aligned with Constructivism pedagogies emphasizing active knowledge construction and learner agency. Research and administrative support (31 studies) covered literature-review assistance, qualitative data-analysis support, course and curriculum development, and drafting of syllabi and policy documents — faculty reported time savings and reduced cognitive load, tempered by concerns about professional identity erosion and deskilling.
Opportunities and Benefits
Six interconnected opportunity domains emerged. GenAI improved learning efficiency and Accessibility through immediate information access and reduced cognitive burden, with students reporting 30–40% time savings on research-intensive assignments; accessibility benefits extended to students with dyslexia, ADHD, and anxiety, aligning with UDL principles. Personalization and Adaptive Learning enabled individualized difficulty, pacing, and instructional strategies at scale — a longstanding challenge in mass higher education. Well-designed applications promoted Metacognition and Self-Regulated Learning, with reflective protocols showing students who documented prompt formulation and output evaluation gained metacognitive awareness. Faculty productivity improved by 25–35% on routine preparation, freeing time for mentoring and innovation. GenAI offered potential democratization of educational resources, particularly for Global South and under-resourced students via 24/7 access, and boosted research and scholarly productivity across the research lifecycle — though researchers emphasized augmentation over automation.
Challenges and Risks
Academic Integrity was the most frequently discussed challenge (89 studies, 71.2%), with 34–37% of students admitting to GenAI use they recognized as potentially violating integrity policies. Conventional formats — essays, problem sets, take-home exams — became vulnerable, forcing reconsideration of what constitutes authentic assessment and original thought. Detection tool limitations proved severe: false-positive rates reached 30–50% for certain writing styles (disproportionately flagging non-native English speakers and neurodivergent students), and simple prompt-engineering strategies reduced detection to near-zero, undermining confidence in surveillance-based responses. Hallucination and accuracy concerns included 10–40% hallucination rates, with medical students detecting AI errors only 44–55% of the time — a patient-safety red flag. Educational Development readiness gaps saw 60–70% of instructors feeling unprepared to integrate GenAI, creating inconsistent student experiences across courses. Finally, Equity and Digital Divide concerns persisted across premium-tool access, algorithmic bias, infrastructure disparities, and AI Literacy gaps correlated with socioeconomic status — stratification that risks amplifying existing educational inequities.
Pedagogical Innovations and Frameworks
Educators responded with several structured innovations. AI Literacy frameworks structured competency development across foundational, conceptual, social, ethical, and emotional dimensions, operationalized through progressive curricula from prompt navigation to policy critique. AI-adapted Technological Pedagogical Content Knowledge (TPACK) (AI-TPACK) distinguished AI-technological knowledge as a specialized dimension intersecting pedagogical and content knowledge; successful integration required balanced development across all three domains. Assessment redesign shifted emphasis from polished products to intellectual processes through process portfolios, oral examinations, personalized context-specific prompts, and transparent AI-integration assignments that require documented usage and critical evaluation. Policy and AI Governance evolved from prohibition toward nuanced, participatory guidance featuring educational rationale, discipline-specific rules, transparency requirements, and student support. Prompt Engineering emerged as both a practical skill and a pedagogical opportunity, with competencies transferable to communication, research-question formulation, and metacognitive planning.
Discussion
Three meta-level findings frame the review's interpretation. First, a curriculum-integration paradox: institutional investment in GenAI infrastructure and Educational Development correlated only weakly (r = 0.12) with actual pedagogical transformation, suggesting technological provision alone is insufficient. Second, a Global North–South divergence: while Global North contexts frame GenAI as an opportunity for Creativity and efficiency, Global South literature emphasizes access, infrastructure, and equity — reinforcing the need for contextually sensitive implementation. Third, an interconnected ethical ecosystem: integrity, privacy, bias, and equity concerns are not independent; detection-focused integrity measures can exacerbate equity problems, and bias-mitigation may limit accessibility. The authors note a maturation trajectory from early 2023 techno-optimism/pessimism toward context-dependent, evidence-based assessments, and flag literature limitations including short time horizons, survey-heavy methodologies, outcome-measurement challenges, and publication bias.
What this means for practice
-
Instructors. Redesign assessment around documented process — portfolios, oral examinations, and transparent AI-use statements — rather than polished products, since Academic Integrity was the most frequently discussed challenge (89 studies, 71.2%) and 34–37% of students admitted use they recognized as potentially violating integrity policy.
-
Instructors. Require students to document prompt formulation and output evaluation; the review links reflective protocols to measurable metacognitive awareness and warns that unexamined use produces over-reliance that inhibits metacognitive development and peer collaboration. Treat AI literacy as a curriculum competency in its own right rather than a one-off workshop.
-
Instructors. Check that AI-adapted AI-TPACK development is balanced across technological, pedagogical, and content knowledge, since the review found integration succeeded only when all three domains advanced together.
-
Administrators. Fund Educational Development at the same priority as tool procurement: 60–70% of instructors felt inadequately prepared to integrate GenAI, and institutional investment correlated only weakly (r = 0.12) with actual pedagogical transformation. The same budget decision applies to integrity: refuse AI-detection tools as the integrity response, since false-positive rates reached 30–50% for certain writing styles and fell disproportionately on non-native English speakers and neurodivergent students, while simple prompting strategies drove detection to near zero.
-
Researchers. Target equity-focused and experimental designs: only 23 of 125 studies (18.4%) came from Global South contexts, and survey research (n = 47, 37.6%) dominated a corpus that needs randomized, validated-outcome studies of whole-institution implementations.
Limitations
- Of the 125 included studies, 37.6% (n = 47) were quantitative surveys, so self-report bias and social-desirability effects carry much of the evidence, while experimental designs with randomization, control conditions, and validated outcome measures remained scarce.
- Coverage is geographically uneven: 23 studies (18.4%) came from Global South contexts across 42 countries, and the headline adoption figure — 92% of students by 2025 — is drawn from UK student data.
- The temporal window is short: most studies capture only immediate or short-term effects after GenAI's November 2022 release, with no multi-year tracking of learning outcomes or career trajectories, and most examine a single course or program rather than institution-wide implementation.
- The authors flag publication bias and measurement problems: the novelty of GenAI favors positive or striking findings, and constructs such as critical thinking and creativity lack consensus measurement, so perceived effectiveness may be inflated.
Citation
Rathnayake, P. B. (2026). Generative AI in Higher Education: A Systematic Review of Opportunities, Challenges, and Pedagogical Innovations (2022–2025). EdArXiv preprint.