Research Article
Educational integrity in GenAI-augmented assessment: making judgment visible
Sharma's paper in the International Journal for Educational Integrity 22:6 (received 18 December 2025, accepted 9 March 2026) argues that educational integrity in GenAI-augmented assessment is better enacted through assessment design than secured through detection. It is a conceptual-theoretical design informed by reflective professional practice, written from Brock University's Department of Educational Studies in St. Catharines, Canada, and it extends Eaton's (2023) postplagiarism by translating its ethical commitments into structural assessment principles — something postplagiarism itself does not address.
The central move installs judgment as the evaluative locus of integrity: the learner's capacity to weigh options, justify academic choices, and assume responsibility in relation to disciplinary standards and ethical expectations (Sadler 1989; Torrance 2007), exercised under epistemic uncertainty rather than as a purely cognitive skill. The paper claims its contribution in three interrelated ways — reframing postplagiarism as a structural challenge for assessment design, theorizing judgment as integrity's evaluative locus, and identifying integrity artifacts through which assessment functions as governance.
Synthesis: Sharma extends Eaton's (2023) postplagiarism from an ethical orientation into assessment design, arguing that detection and verification-based approaches to Academic Integrity are misaligned with GenAI-augmented work, where human judgment and technological generation are deeply entangled and resist clean separation. Integrity is reframed as a pedagogical practice enacted through Evaluative Judgment: the capacity to weigh options, justify choices and assume responsibility under epistemic uncertainty. The paper illustrates integrity artifacts — annotated decision trails, verification and accountability, oral defense, and draft differences — and positions detection as supplementary rather than foundational infrastructure. It is theoretical, drawn from the author's professional teaching context, and offered to invite reconsideration of integrity frameworks rather than to generalize across institutions.
Key Findings
- Detection rests on an untenable assumption. Generative text is probabilistic, non-deterministic, and shaped through iterative human prompting and revision, so definitive attribution to human or machine sources is unreliable (Walton et al. 2025), and isolating a stable "signal" of technological authorship struggles to distinguish unethical delegation from responsible engagement.
- Detection asks the wrong question. It asks whether GenAI was used rather than how decisions were made. Students who reject inaccurate outputs, revise drafts, and align content with disciplinary criteria (Walton et al. 2025) produce evidence detection renders invisible, and instructors "may be compelled to interpret probabilistic outputs without shared evidentiary standards, shifting their role toward investigation rather than assessment" (Lynch et al. 2021).
- Detection is recalibrated, not abolished. Formal adjudication still needs mechanisms for clear misrepresentation or deliberate outsourcing of intellectual labor, so detection can serve as one layer within a broader, multi-layer AI Governance framework — but not as the primary infrastructure of integrity.
- Assessment is an integrity apparatus. It signals what counts as evidence, how responsibility is interpreted, and which behaviors are valued (Sadler 1989; Torrance 2007). The distinction from Authentic Assessment is epistemic: authenticity asks whether a task mirrors real-world practice, whereas integrity-oriented design asks whether learners can justify decisions against disciplinary standards, making integrity an explicit, assessable criterion inside the task.
- Judgment differs from reflection and metacognition. Reflection considers thinking retrospectively (Schön 1983) and Metacognition regulates cognitive processes (Zimmerman 2002); judgment is evaluative (Sadler 1989), sitting at the intersection of cognition, ethics, and professional accountability.
- Responsibility is redistributed across systems. Institutions recognize GenAI-augmented authorship as normative, educators design assessments making ethical decision-making visible, and learners articulate judgment against disciplinary criteria, so integrity becomes a property of educational systems rather than of individual behavior.
Why detection loses its footing
Generative systems act as interactive collaborators responsive to prompting and revision (Holmes et al. 2022), producing work that resists clean separation into "human" and "machine" contributions; efforts to preserve pre-GenAI frameworks through stricter enforcement risk misrecognizing that hybrid production (Walton et al. 2025), making attribution unreliable as evidence.
Pedagogically, detection models prioritize identifying inputs over evaluating judgment, and uniform detection standards risk inconsistency as GenAI use becomes context-dependent (Fengchun and Cukurova 2024). Ethically, surveillance-based regimes risk eroding Trust by positioning students as potential violators rather than ethical agents. Sharma does not overreach: procedural defensibility and regulatory accountability still require mechanisms supporting formal adjudication, so detection survives as one layer. The argument is for recalibration — governance must prioritize designs evaluating transparent, justified decision-making, with detection supplementary rather than foundational.
Judgment, artifacts, and the burden of interpretation
Assessment becomes the primary site where integrity is enacted, because what a task treats as evidence teaches learners what responsibility means. The integrity artifacts are illustrative rather than evidentiary, developed in the author's professional teaching context and selected for representing the paper's theoretical commitments. As reflexive illustration, Appendix A reproduces a decision trail from the manuscript's own production: the author asked a GenAI tool to role-play an editor and propose revisions, then recorded accepted suggestions (clarifying the contribution beyond Eaton 2023, defining judgment early, distinguishing the examples from authentic assessment, shortening the title from 17 words) alongside rejected ones.
Artifacts are explicitly not proofs. Requiring learners to document reasoning may privilege those more fluent in reflective discourse, risking "the replacement of one compliance regime with another"; artifacts need contextual interpretation, shared criteria, and professional calibration, since judgment as evidence of integrity "remains relational and situated rather than mechanically verifiable."
Equity, trust and institutional design
Institutional responses emphasizing prohibition, detection, and enforcement treat integrity as a problem managed through surveillance or policy compliance, which the paper argues risks undermining trust, equity, and pedagogical coherence. Detection tools operate probabilistically and are often misinterpreted as definitive evidence rather than indicators requiring contextual judgment (Walton et al. 2025), and surveillance regimes erode trust particularly when expectations for acceptable use remain ambiguous or inconsistently applied (Bretag 2019; Selwyn 2019).
The equity argument is load-bearing. Learners engage generative tools from unequal positions shaped by language proficiency, prior opportunity, disciplinary familiarity, and access to institutional support (Agheorghiesei and Bercu 2022), and for some GenAI functions as linguistic or cognitive Scaffolding rather than substitution. Detection-centered models may therefore disproportionately disadvantage those learners (Francis et al. 2025), and where integrity is inferred from surface textual features rather than articulated reasoning, inequities are more likely to be penalized than addressed. Because norms of authorship are historically and culturally situated and plagiarism is interpreted differently across systems (Abasi and Graves 2008; Bretag 2016), integrity requires explicit dialogue rather than universal procedural rules.
What this means for practice
- Instructors. Assess reasoning rather than reading detection output as evidence: use annotated decision trails, verification logs, oral defense, and learner-selected draft comparisons, scoring quality of justification rather than volume of reflection. Focus annotation on selected decision points, and in large classes use brief structured conferences, small-group dialogic checks, or selective sampling with shared calibration, so dialogic formats evaluate conceptual understanding rather than oral fluency.
- Assessment designers. Build integrity into task architecture. Verify reasoning by triangulating evidence across paired tasks, connecting to designs that enhance assessment validity in AI-vulnerable contexts (Roe et al. 2025), and treat GenAI inaccuracies as objects of verification responsibility rather than automatic evidence of misconduct.
- Administrators. Recognize GenAI-augmented authorship as normative in policy, fund professional learning so markers interpret articulated judgment consistently, and treat integrity as a property of the system rather than a trait of individual students; designs inviting articulation of responsibility signal institutional confidence in learners.
Limitations
- No empirical findings and no causal claims. The practices come from the author's professional teaching experience and are offered illustratively, so the argument is theoretical; data availability is stated as not applicable, as no datasets were generated or analyzed.
- A single supportive setting. Reflective engagement in one higher education context enables grounded insight but introduces possible confirmation bias and contextual limitation, and the practices may not transfer to differently resourced or less supportive environments.
- Interpretive reliability is unresolved. Judgment is situated and its articulation varies across learners, so designs requiring justification depend on educators' capacity to interpret reasoning consistently — an open question of evaluative coherence the paper refers to future research.
- Visible judgment is not assured ethical practice. The practices may become performative, with learners aligning discourse to expectations without internalizing commitments; making judgment visible creates conditions for ethical dialogue without guaranteeing it.
Connected Concepts
- Academic Integrity — The paper's subject: integrity reframed as pedagogical practice rather than compliance
- Evaluative Judgment — The construct the argument installs as the evaluative locus of integrity
- Assessment — Positioned as the integrity apparatus through which judgment becomes visible
- Assessment Validity — The design goal served by paired verification tasks
- Authentic Assessment — The adjacent movement the paper distinguishes its designs from
- AI Detection — Recalibrated from primary infrastructure to one layer of governance
- AI Use and Disclosure Statements — What annotated decision trails ask learners to articulate
- Generative AI — The technology whose outputs and errors become objects of verification responsibility
- Metacognition — Distinguished from judgment in the paper's conceptual work
- Self-Regulated Learning — Formative assessment and criteria transparency as the mechanism behind verification practices
- Trust — What surveillance-oriented regimes erode
- Equity — Detection-based inference penalizes learners who use GenAI for linguistic or cognitive support
- AI Governance — Institutional policy reoriented from enforcement toward assessment architecture
- Higher Education — The sector the argument addresses
Connected Articles
- How university students work on assessment tasks with generative AI: matters of judgement — The study of students' judgment in GenAI assessment tasks that underpins Sharma's argument
- Assessment twins: An approach for strengthening assessment validity in the age of generative AI — Paired-task designs that verify learning outcomes in AI-vulnerable summative assessment
- The Impact of Generative AI on Academic Integrity of Authentic Assessments Within a Higher Education Context — Evidence that authentic assessment alone does not safeguard integrity
- The Integrity of Psychology Assessments in the AI Age: A Critical Examination — Program-level evidence that the pass boundary and detection fail in different ways
- Beyond Detection: Redesigning Authentic Assessment in an AI-Mediated World — A design-side companion argument for moving past detection
- Generative AI as a Design Variable: An Evidence-Centered Framework for Principled Governance in STEM Assessment — Treating GenAI as a design variable in assessment governance
- Cheating or not cheating? Rethinking AI-giarism and academic integrity through secondary students' ethical reasoning — Students' own ethical reasoning about AI use as the object of integrity work
- "Should I Tell My Teacher?" Student AI Disclosure Practices, Stigma, and Self-Regulated Learning in Higher Education — What disclosure practices look like when the burden falls on learners
Citation
Sharma, S. (2026). Educational integrity in GenAI-augmented assessment: making judgment visible. International Journal for Educational Integrity, 22(6).