On this page

Synthesis: Kofinas, Tsay, and Pike (2025) ran a series of experiments across two UK universities in which experienced academic markers judged human-authored, GenAI-modified, and GenAI-generated assessments. They found that markers generally could not distinguish assessments with GenAI input from those without, that the presence of GenAI nonetheless affected how markers approached marking, and that the level of authenticity in an assessment had no impact on the ability to safeguard against or detect GenAI use. The authors conclude that current assessment approaches are susceptible to GenAI manipulation and that authentic assessments alone cannot protect academic integrity; the focus must shift to assessment design favouring process-based, performative, and synchronous interpersonal forms.

Key Findings

  1. Markers cannot reliably detect GenAI. Markers generally could not distinguish assessments that had GenAI input from those that did not, even when the assessment was authentic.
  2. GenAI alters the marking process. The mere possibility that GenAI was used (and suspicion of tampering) changed markers' behaviour and judgement, producing both false positives and false negatives—so students' learning is not assessed correctly.
  3. Authenticity offers no safeguarding benefit. The level of authenticity in an assessment had very limited impact on the ease of detecting GenAI usage; AI could generate authentic assessments that passed faculty scrutiny.
  4. Consequentialist reasoning drives misuse. Students weigh the benefits (efficiency, ease) against risks of getting caught; when benefits outweigh risks, they may use GenAI in ways that compromise integrity.
  5. GenAI assessment is cheap and 24/7. Unlike essay mills or contract cheating, GenAI can produce high-quality, seemingly original work for free or very cheaply, undermining authorship detection and circumventing the effort required to learn.
  6. Output-based written assessments are vulnerable. Common written formats—reports, essays, and take-home exams—are susceptible, whereas performative, synchronous, and socially experiential assessment formats are more resistant.

Implications

  • Authentic assessments are not a panacea for protecting academic integrity from GenAI; policy and practice must keep the focus on assessment design rather than the format label alone.
  • Assessments of learning should shift from assessing output to focusing on process and workplace relevance, a paradigmatic shift from written to synchronous interpersonal (performative, dialogic, oral) assessment models.
  • This move has far-reaching consequences for the academy: if written assessments can no longer be trusted as a reliable indicator of learning, marking, moderation, and credentialing practices must be reconsidered.
  • The findings underpin the argument for authentic, presence-required assessment formats that value process and social experiential learning, complementing the design-for-presence agenda in agentic AI debates.

Connected Concepts

Connected Articles

Citation

Kofinas, A. K., Tsay, C. H.-H., & Pike, D. (2025). The impact of generative AI on academic integrity of authentic assessments within a higher education context. British Journal of Educational Technology, 56(6).