On this page

Synthesis: Balalle and Pannilage (2025) present a PRISMA-based systematic literature review (25 studies from 1,443 records across Scopus, PubMed, DOAJ, and BASE) examining the impact of artificial intelligence on academic integrity in higher education. The review finds a genuine research gap — AI and academic integrity sit in the same keyword cluster but few studies cover both — and documents that AI functions as both a threat to integrity (AI-generated writing, paraphrasing tools, contract cheating) and a tool for detection (Turnitin AI/similarity scoring). Its central call is for institutions to build a culture of academic integrity through clear policy, assessment redesign, and ethics training, rather than relying on detection software alone.

Methods

The review used a PICO-framed research question ("What is the role of AI in influencing academic integrity, and how can educational institutions ensure ethical AI usage?") and the PRISMA 2020 flow for study selection:

  • 1,443 records identified (PubMed 62, DOAJ 1,136, Scopus 235, BASE 7, plus 3 expert recommendations); 32 duplicates removed → 1,408 screened → 78 full-text retrieved → 25 included in the final analysis.
  • Risk of bias was assessed with the Cochrane ROBINS-I tool via Nested Knowledge's semi-automated platform. The most significant bias was in participant selection; 9 studies showed selection concerns, and several others showed bias due to confounding, missing data, or selective reporting.
  • A VOSviewer keyword co-occurrence network (4 clusters) surfaced the field's structure: cluster 1 (academic integrity, academic misconduct, AI, student character), cluster 2 (academic writing, generative AI, large language model), cluster 3 (ChatGPT, higher education, quality assurance), cluster 4 (plagiarism).
  • The most-cited works were Cotton et al. (2024) "Chatting and cheating" (755 citations), Sullivan (2023) "ChatGPT in higher education" (221), and Crawford et al. (2023) "Leadership is needed for ethical ChatGPT" (188).

Key findings

AI as both threat and detection tool

The review documents AI's dual role:

  • As a threat: Students use generative AI (ChatGPT, Google Bard, Writesonic, Jasper, Wordtune, automated paraphrasing tools) to complete assignments, reducing originality and risking the credibility of qualifications. Non-native English speakers show a high tendency to breach integrity when struggling to write in English.
  • As a detection tool: Most institutions use Turnitin (similarity + AI content rate) integrated into learning management systems; online proctoring systems and cameras are used for exam monitoring. However, the review cautions that a low AI content rate may be a false positive, and that ai-detection tools are unreliable for AI-generated work — institutions should not rely on them alone.

Detection is an industry, not a solution

Academic cheating and detection have become "thriving industries" — companies sell both AI-writing/humanizing software and detection software. The review argues institutions can develop their own manual detection procedures and train lecture panels, and that multiple assessment methods (oral exams, traditional writing tests, essays) should be used to detect AI misconduct rather than software alone (citing Bozkurt 2024).

Culture of integrity over policing

Across the included studies, the dominant recommendation is preventive and cultural rather than purely technological:

  • Clearly define what constitutes academic misconduct (plagiarism, AI writing, cheating) and establish explicit policies and expectations.
  • Develop an honor code agreed by students and faculty, and a jointly-developed campus AI-use policy, formally documented and distributed.
  • Build a culture of academic integrity from orientation onward, with regular workshops and training on citation practices, academic ethics, plagiarism, and AI detection tools.
  • Redefine academic integrity for the technology-based era — the old definition no longer fits.
  • Balance detection with assessment redesign (some institutions have returned to pen-and-paper exams), while recognizing that banning AI is difficult to enforce (ChatGPT is hard to prevent, per Sharples 2022).

AI also has legitimate value

The review is balanced: AI can enhance writing efficiency, improve non-native English writing, act as a virtual tutor students ask questions to without hesitation, and improve learning abilities (Darvishi et al. 2024; Maphoto et al. 2024; Milano et al. 2023). The task is to harness these benefits while upholding ethical standards, not to ban the tools.

What this means for practice

  • Administrators. Build a culture of academic integrity rather than an enforcement regime: the review concludes that institutions need balance between preventing misconduct and preserving academic freedom and innovation, with expectations agreed and published rather than imposed.
  • Administrators. Publish explicit expectations for acknowledging AI use in academic activities, since the reviewed studies report that staff and students want those policies in place and the field has had to construct shared definitions of integrity for AI-era work.
  • Instructors. Do not let a tool score decide a misconduct case. Institutions can integrate Turnitin directly into the LMS for detection, but in-person exams monitored by invigilators — or proctored online settings using cameras — remain the configurations in which staff can directly observe student behavior, so pair detector output with observed conditions.
  • Instructors. Treat writing difficulty as a risk factor: the review found non-native English-speaking students showed a high tendency to breach integrity when struggling to write in English, so pair integrity expectations with concrete academic-writing support.
  • Administrators. Commit to continuous monitoring, evaluation, and adaptation of teaching and assessment practices, as the reviewed institutional studies recommend, instead of a one-off policy launch.

Limitations

  • The evidence base is small and heterogeneous: 25 studies were included from 1,443 records, the total participant pool across them is 2,134, and the corpus mixes quantitative studies with review papers.
  • The risk-of-bias assessment flagged participant selection as the most significant problem — nine studies fell in that domain, including Maphoto et al. (2024), which sampled 70 participants from a population of 14,000 — and four studies raised concerns about selective reporting of results.
  • The database search was run across the entire publication period with no year filter, 11 of 89 full-text reports could not be retrieved, and several included entries report no information about their participants.
  • The evidence is recent and fast-moving: the most-cited sources are 2023–2024 responses to ChatGPT's November 2022 release, so the findings may not extend to later tools and policies.

Citation

Balalle, H., & Pannilage, S. (2025). Reassessing academic integrity in the age of AI: A systematic literature review on AI and academic integrity. Social Sciences & Humanities Open, 11, 101299.

Embed this page

Copy the code below to embed a chromeless version of this page in a learning management system or other website. The embedded view hides the site header, navigation, and footer.