On this page

Synthesis: This PRISMA-guided scoping review maps 38 empirical studies of what higher education students themselves say about Academic Integrity when they use generative AI. Ji reports a field that appeared in 2023 and exploded in 2025, synthesising student voices into four themes: conceptual, policy and pedagogical ambiguity; ethical agency; the gap between ethical awareness and practice; and diversity by gender, level, discipline and culture. The review's central claim is that students are neither passive recipients of GenAI nor culprits without moral constraint, but active agents navigating an ethical grey zone, so institutional responses should be co-created, educative and context-sensitive rather than punitive.

Key Findings

  1. 38 studies, four themes, three years. The review screened 142 documents down to 38 included studies (34 from the database search plus four from reference lists), analysed with Braun and Clarke's six-phase thematic analysis. Two studies appeared in 2023, eight in 2024, 22 in 2025 (58% of the corpus) and six in 2026, across 28 journals, most often the Journal of Academic Ethics (n = 5). Only about a third of studies report formal funding; 22 of 38 (58%) declare none.
  2. The evidence is survey-heavy and mostly measured with home-made instruments. Eighteen studies (47%) were quantitative, 12 (32%) qualitative and eight (21%) mixed methods. A questionnaire was the data-collection tool in 34 studies, and 29 of 38 (76%) developed their own instrument, which Ji reads as validated scales being perceived as insufficient for GenAI-era questions. Rigour is uneven but improving: nine studies report pilot testing, eight expert review and seven reliability coefficients. Samples range from Ofem et al.'s (2025) more than 4,500 Nigerian students to Dillion et al.'s (2024) three participants, and 27 of the 38 included postgraduates.
  3. Theme 1 is ambiguity. Generating an entire task is read as cheating; grammar checking and brainstorming are largely acceptable; and the vast middle ground of paraphrasing AI text, generating ideas or outlines, and translating or writing parts of a task forms an ethical grey area. Ji adopts Chan's (2025) term "AI-giarism" for the emergent misconduct this creates, and notes that Ecuadorian students in Nelson et al. (2025) treated chatbot-generated text as dishonest while judging machine translation ambivalently. Calls for clear guidelines, ethics training and dialogue recur across Indonesian (Barus et al., 2025), South Asian (Hossain et al., 2026) and 70-plus-country samples (Yusuf et al., 2024), and Brickhill et al. (2025, p. 971) record students' anxiety about unintentional misconduct alongside their request for more conversation.
  4. Theme 2 is students' ethical agency, in four forms. First, oversight-preserving collaboration: Canadian graduate students stressed "the importance of human agency and decision-making in the AI-assisted research process, and the need for critical evaluation and personalisation of AI-generated content to maintain authorship" (Sabbaghan & Eaton, 2025, p. 1860). Second, de facto policymaking, because the guidance vacuum leaves students building personal rules (Barus et al., 2025) in what Tsao (2025) calls invisible labour, and reading silence as "a silent approval of their actions" (Giray et al., 2026). Third, risk calibration: Croatian students (Črček & Patekar, 2023) used ChatGPT freely for ideas but cautiously for writing. Fourth, restraint: Nigerian students limited use for fear of crossing ethical lines (Nwagbara, 2025). Bearman et al.'s (2026) "coming from me" imperative names the conviction underneath, and Ecuadorian EFL students worried more about over-reliance stunting their writing than about being caught.
  5. Theme 3 is the gap between awareness and practice. Croatian students who judged AI-written assignment sections unethical sometimes wrote them that way anyway (Črček & Patekar, 2023), and Huang et al.'s (2025) "Ethical Dissonance Index" sorted 522 Chinese students into four clusters, one of which frequently did what it viewed as illegitimate. Ofem et al.'s (2025, p. 159) structural equation modelling of 4,679 Nigerian students found that "students with positive perceptions of ChatGPT were more prone to use it for dishonest academic purposes," with positive integrity attitudes a "significant negative mediator." Rationalisation does the work: Filipino undergraduates recast GenAI cheating as "a legitimate academic survival strategy rather than misconduct," evading detection in courses they saw as irrelevant to their careers, with peers normalising the practice (Giray et al., 2026). Workload, grade pressure, stress, low confidence and, for multilingual students, linguistic disadvantage also serve as justifications (Hysaj et al., 2025; Nelson et al., 2025).
  6. Theme 4 is diversity, and it cuts against generalisation. Gender findings conflict: male students reported more confidence and female students more concern (Bikanga Ada, 2024; Prohorovs et al., 2025), male Chinese students were likelier to report plagiarism while female students rated its harm as more severe (Huang et al., 2025), yet Gruenhagen et al. (2024) found no gender difference. US graduate students judged AI use more strictly than undergraduates (Lund et al., 2025b), and "science students... more likely to engage in online plagiarism than... humanities students" (Huang et al., 2025, p. 1). Culture is the widest axis: under identical policies, Canadian students were consistently likelier than Korean students to call GenAI use unethical (Harrington et al., 2025).

How the review was conducted

Ji searched Web of Science, Scopus and ERIC between October and late November 2025, combining integrity, GenAI, student and perception search terms. The search returned 142 documents, including 10 conference papers. Screening excluded 108 of them (45 duplicates, 63 that failed the study-type, topic, participant or language criteria), and four further papers came from reference lists. Inclusion required empirical peer-reviewed English-language journal articles with higher education students as participants and student integrity voice as a research focus. Extraction ran in two phases: a self-developed coding table for bibliometric and methodological data, then thematic analysis via Braun and Clarke's six phases. Ji is explicit that every reviewed study rests on self-report, and that students may under-report unethical use or overstate concern when integrity is the stated research focus.

How widespread use is, and how students describe it

Several studies report that over 95% of participating students had used GenAI (Barus et al., 2025; He et al., 2026); the lowest was about a third (Kazley et al., 2025). Students use the tools mainly for searching information and improving comprehension (Chan, 2025; Gruenhagen et al., 2024; Sajja et al., 2025; Yusuf et al., 2024), and for writing-related language support: paraphrasing, brainstorming, grammar and structure, and translation (Črček & Patekar, 2023; Hysaj et al., 2025; Nwagbara, 2025). Almost all reviewed studies report students crediting GenAI with efficiency and time savings (Bearman et al., 2026; Tsao, 2025) and treating it as a personalised aid where teacher support or materials are scarce (Fajt & Schiller, 2025; Sajja et al., 2025). Ji concludes that students view ChatGPT as a legitimate learning tool rather than a cheating aid, citing Perdana et al.'s (2026) "augmented Metacognition" for extending rather than replacing thought. The caveat is social desirability: with integrity as the research topic, admissions of graded-assignment use may be suppressed.

What the reviewed studies recommend

The implications converge in three directions. First, from retrospective detection to proactive reasoning: nearly all studies ask for a shift away from rule-based enforcement and detection toward educative approaches, with Giray et al. (2026, p. 18) arguing that "traditional punitive approaches to academic integrity are insufficient in the AI era" and that institutions need "proactive ethical education, AI literacy programs, and policy frameworks that acknowledge technological realities." Ji adds that AI ethics and AI-giarism belong in the curriculum rather than in one-off literacy sessions (Gruenhagen et al., 2024), and that ownership and moral reasoning motivate students more than fear of punishment (Bearman et al., 2026).

Second, from ambiguity to co-created clarity: because unclear guidance is read as tacit permission, studies want detailed, accessible guidelines with concrete examples of permissible and impermissible use and explicit differentiation between assistive tools such as Grammarly and generative tools such as ChatGPT (Lund et al., 2025a, 2025b; Tsao, 2025; Yusuf et al., 2024), co-created by leaders, faculty and students (Shaukat et al., 2025).

Third, from universal principles to contextual sensitivity: assessment design is named as a major driver of breach, because tasks students see as irrelevant or overly burdensome get reframed as obstacles to survive (Črček & Patekar, 2023; Hysaj et al., 2025; Nelson et al., 2025). The studies call for tasks GenAI cannot do well, such as fieldwork, oral exams, contextual essays and case-based analysis (Tsao, 2025; Barrett & Pack, 2023), and for personalised reflective writing in which the student's own context carries the work (Dillon et al., 2024). Since norms differ by discipline, level and culture (Kazley et al., 2025; Harrington et al., 2025), guidelines should be differentiated and faculty supported through professional development to design discipline-specific valid tasks (Kamoun et al., 2024; Sajja et al., 2025; Synekop et al., 2024).

Gaps, limits and what the review itself cannot do

The reviewed studies name four recurring weaknesses. Samples are context-specific and rarely cross-national, so generalisability is limited, and coverage skews toward Western, English-speaking, high-income systems despite emerging Global South research (Nelson et al., 2025); sizes are often small. All depend on self-reported survey and interview data, open to social desirability bias, and several call for more diversified sources to triangulate findings. Most consequentially, every reviewed study is cross-sectional: the field lacks longitudinal work tracking students as tools and policies change, and lacks intervention studies testing whether ethics curricula, redesigned assessments or different policy communications work. The studies also suggest looking beyond writing, to coding, data science and music composition. Ji adds three limits of the review itself: three databases only, English-language publications only, and a technology moving too fast for any such map to stay current.

Connected Concepts

Connected Articles

Citation

Ji, Y. (2026). Academic integrity in the age of generative AI: A scoping review of research on higher education student voices. Journal of Academic Ethics, 24, 78.

Connected FAQs

Embed this page

Copy the code below to embed a chromeless version of this page in a learning management system or other website. The embedded view hides the site header, navigation, and footer.