On this page

Synthesis: Munoz and colleagues coded every generative AI-related academic misconduct case at one regional Australian university over three years, 1,162 cases carrying 1,855 discrete evidence items, then classified each item with an empirically derived 15-code taxonomy and rated it for probative value on three dimensions drawn from legal evidence scholarship: relevance, credibility, and inferential force. The strongest evidence types were the ones that do not depend on probabilistic classification of text, namely student admissions, observed prohibited exam behaviour, and independently verified fabricated references, while AI-detector outputs attracted the weakest ratings of any category. The paper's more uncomfortable result is structural: evidence quality showed no reliable relationship to case outcomes, because institutional procedures set no minimum evidentiary threshold and impose no requirement to weigh probative value before a case is progressed. The authors offer the taxonomy and a credentials framework as a diagnostic for codifying evidentiary standards in Academic Integrity policy.

Key Findings

  1. GenAI cases grew from 11.9% to 36.9% of all misconduct allegations while total volume stayed flat. Across 5,127 case records lodged between January 2023 and December 2025, GenAI-related allegations rose from 11.9% of misconduct cases in 2023 to 20.0% in 2024 and 36.9% in 2025 (χ²(2) = 320.93, p < .001, Cramér's V = 0.250, approaching a moderate effect). The 1,162 GenAI cases are 22.7% of the institutional misconduct corpus, which the authors read as GenAI allegations displacing other allegation types rather than adding cases.
  2. Fifty-eight cases (5.0%) rested on assertion alone. These records held no identifiable evidence item and received only the Assertion Only code, so 63 items (3.4% of the corpus) were allegations with no supporting evidentiary content. The share of items coded Assertion Only fell from 7.6% in 2023 to roughly 2.5% by 2025, which the authors read as more consistent documentation of discrete evidence.
  3. Fabricated references became the largest structural shift in investigative practice. Non-existent or fabricated references accounted for 414 items (22.3%) across 356 cases (30.6%), rising from 10.4% of coded items in 2023 to 30.0% in 2025. They carried high relevance (99.2%) and credibility (91.6%) but predominantly moderate inferential force (98.0%) by design: a student may fabricate independently, authorised or incidental AI use can introduce a fabricated reference, and reference hallucination leaves no signal in work that cites nothing.
  4. The strongest evidence is generated by the investigation, not the allegation. Oral interview admissions were the most frequent substantive code (410 items, 22.1%; 399 cases, 34.3%) and 97.2% were rated Strong inferential force. Written admissions (10 items) were 100% Strong, proctoring records 86.1% Strong, and observed prohibited exam behaviour 68.4% Strong. Because admissions emerge only after the institution has decided the allegation is sufficient to proceed, the strongest evidence in many cases arrived late.
  5. Detection and textual evidence were weak, and detector output then vanished. Turnitin or similarity reports appeared in 73 items (3.9%) with 100% Weak inferential force, and standalone detector outputs such as GPTZero in 69 items (3.7%) with 100% Low credibility. Detector evidence fell from 5.8% of items in 2023 and 8.8% in 2024 to 0.5% in 2025 (n = 5), consistent with institutions recognising that it cannot discharge the civil standard of proof. Textual signals were frequent but inconclusive: AI-typical content patterns were the second-largest code (302 items, 16.3%; 239 cases) yet 70.6% received Weak inferential force, and comparative style or level anomalies split almost evenly between moderate (51.4%) and weak (48.6%).
  6. Evidence quality did not predict outcomes, which the authors call a finding about the system. Associations were significant but modest at the bivariate level, for example cases with more evidence items at the Academic Integrity Officer stage were less likely to be dismissed (OR 0.31, 95% CI 0.06–0.89). Escalation was governed by policy triggers, and the misconduct-versus-poor-academic-practice distinction by first-offence status and deliberateness rather than evidentiary credentials. "There is no requirement for investigators to assess the probative quality of evidence before progressing an allegation, no minimum evidentiary threshold at any stage of the pipeline."

How the study was conducted

The setting was a single regional Australian university registered with TEQSA, using an in-house case management system in place since 2017. Allegations move through a staged pipeline: teaching staff review the evidence and interview the student, an Academic Integrity Officer conducts an independent review and may escalate, and a Faculty Investigation Committee holds a hearing and determines the outcome. At every stage the standard is the balance of probabilities, with the burden of proof on the institution. The median interval from allegation to finalisation was 12 days (interquartile range 6–20), so students were typically asked to account for their work close in time to submission.

GenAI cases were extracted in three passes: a categorical filter on misconduct type (the GenAI label was introduced mid-2023), a keyword search of the two mandatory free-text fields using three signal groups covering content-generating tools, text-modifying tools and fabricated-citation terms, and a deterministic five-rule sequence for ambiguous cases, with 12 cases reviewed by hand. Coding used framework analysis following Gale et al. (2013), producing a hybrid deductive-inductive taxonomy constrained by the Anderson et al. (2005) credentials framework and derived from the corpus under three criteria: internal coherence, distinctiveness and recurrence. The final taxonomy holds 15 codes in eight substantive categories.

Each item was rated for relevance and credibility on High/Moderate/Low, and for inferential force, what an item establishes taken alone, on Strong/Moderate/Weak. Coding was a two-stage hybrid of rule-based pattern matching and a locally hosted Qwen 2.5:14b configured to produce identical outputs for identical inputs, with all case data processed on institutional infrastructure. Reliability was established between humans first: four co-investigators independently coded twelve cases, reaching 94.6% majority agreement (Fleiss' κ = 0.619; mean pairwise Cohen's κ = 0.625). Ten refinement cycles raised those to κ = 0.646, Cohen's κ = 0.651 and 100% agreement, and generalisation was tested on a held-out sample of 54 cases and 93 evidence items coded blind, where human–human agreement (AC1 = 0.896) and human–LLM agreement (AC1 = 0.918) met the non-inferiority criterion (Δ = −0.02, 95% CI −0.04 to −0.00).

What the evidence is worth, and why that keeps shifting

The central pattern is a misalignment between the evidence types commonly used at the point of allegation and those carrying the strongest inferential force. Evidence resting on probabilistic classification of text clustered at weak to moderate force despite dominating allegation narratives, which the authors attribute to those signals being common in misconduct proceedings generally rather than exclusive to GenAI misuse. The best-performing types are recoverable differently: admissions address the misuse question directly, observed exam behaviour is documented by objective system records, and fabricated references are independently verifiable against external databases. Reference-based evidence still attracts moderate rather than strong inferential force once calibrated against whether it can discharge the standard alone.

The authors also treat probative value as temporal rather than fixed. The force of any GenAI-specific signal depends on the state of LLM capability when the submission was made, so the same evidence type carries different weight over time, and any framework here will need regular recalibration. The rise of fabricated references illustrates the risk: LLMs are improving at generating plausible, retrievable citations at a documented rate, so a signal whose force rests on the unreliability of LLM citation generation will weaken as that generation improves. Institutional practice is converging on an evidence type whose discriminating power is being eroded by the same development driving the caseload.

Implications for policy and practice

The paper's practical claim is that the credentials framework should be applied at the point of allegation, not only at the point of determination. Asking what an item establishes on its own, using the taxonomy to classify evidence and the credentials to weigh it, would bring investigative practice closer to the evidentiary demands of the process it initiates. The authors stress that individual application is necessary but insufficient while evidentiary standards remain uncodified, since inconsistency in evidentiary values is a structural feature of the current system that good judgement cannot overcome at scale.

Two consequences follow. A process that relies on individual expertise for its evidentiary rigour will produce outcomes that vary with that expertise, and allegations raised by less experienced staff increase the risk of false positives. That variation matters for students, who may be pushed into appeal mechanisms when the evidentiary basis is weak, and for institutions, whose findings become hard to defend as consistent and fair. Codifying standards would make the evidentiary basis of proceedings transparent and auditable, which the authors distinguish from mechanically determining outcomes from evidence scores. The paper also documents how policy caught up with practice: the Assessment policy in force through 2024 said nothing about AI or detection software, and from 1 January 2025 a revised policy permitted authenticity checks through approved software while prohibiting the upload of student work to third-party tools, including AI-detection software.

Limitations

The analysis is retrospective, and coded items represent what was documented in case files rather than the full evidentiary basis for decisions; assessors may have relied on tacit reasoning that never entered the written record, and low-frequency codes may be underreported. The calibration sample of twelve cases is small, several codes remain rare even in the held-out sample so per-code estimates are imprecise, and the study was not disaggregated by cohort or discipline because the de-identified dataset lacks those fields and reporting at that level fell outside the ethics approval. The authors also note their positionality as employees of the institution whose data was analysed, mitigated by anonymisation. Finally, the data come from one Australian university; multi-institution replication is identified as the priority for future research.

Connected Concepts

  • Academic Integrity — the paper's subject: what evidence institutions hold when they allege GenAI misuse
  • Assessment Validity — probative value as the question of whether evidence bears on the claim at issue
  • AI Detection — detector output is the weakest-rated evidence type and has nearly vanished from case files
  • Evaluative Judgment — panels weigh evidence through unstructured professional judgement
  • Educational Measurement — the taxonomy's rating scales and inter-coder reliability work
  • Remote Proctoring — proctoring records and observed prohibited exam behaviour are among the few Strong evidence types
  • Educational AI Policy — the case for codifying evidentiary standards in misconduct procedures
  • AI Governance — staged investigation, the balance of probabilities, and the burden on the institution
  • Student Experience — weak evidentiary bases push students toward appeals
  • Higher Education — institutional practice across a three-year period of GenAI caseload growth
  • Trust — defensibility of findings depends on explicit, auditable evidentiary standards
  • Hallucination Risk — fabricated references are the most frequent concrete GenAI signal
  • AI Misuse and Learning Harm — poor academic practice as an educational rather than disciplinary outcome
  • Reducing AI Misuse — what investigators can actually establish when they suspect misuse

Connected Articles

Citation

Munoz, A., Hinchcliff, M., Langfield, C., & Rogerson, A. (2026). How strong is the evidence in generative AI-related academic misconduct allegations? A mixed-methods analysis. International Journal for Educational Integrity, 22(26).

Embed this page

Copy the code below to embed a chromeless version of this page in a learning management system or other website. The embedded view hides the site header, navigation, and footer.