On this page

Synthesis: Sharma's conceptual and practice-informed paper extends Eaton's (2023) postplagiarism framework from an ethical orientation into assessment design, arguing that detection and verification-based approaches to Academic Integrity are misaligned with GenAI-augmented work, where human judgement and machine generation are entangled and resist clean separation. The paper reframes educational integrity as a pedagogical practice enacted through Evaluative Judgment: the learner's capacity to weigh options, justify academic choices and assume responsibility under epistemic uncertainty. It identifies four illustrative assessment practices as integrity artefacts, annotated decision trails, verification and accountability practices, oral defence and dialogic accountability, and draft differences with version history, and positions detection as a supplementary rather than foundational layer of integrity infrastructure. The argument is theoretical, drawn from the author's professional teaching context, and is offered to invite reconsideration of integrity frameworks rather than to generalise across institutions.

Core Argument

  1. Integrity frameworks were built for a lone author. GenAI has unsettled assumptions about authorship, originality and assessment, and institutional responses organised around detection technologies, usage bans and revised policies work to preserve pre-GenAI models of independent authorship. Such responses risk prioritising compliance over learning and surveillance over trust (Bretag 2019; Selwyn 2019). Sharma's framing is that debates about GenAI in higher education are "fundamentally debates about educational integrity," and that the consequential question is not whether GenAI was used but how responsibility and ethical judgement are defined and enacted when academic work is produced through hybrid human-GenAI systems.
  2. Detection cannot resolve hybrid authorship. Generative text is probabilistic, non-deterministic and shaped through iterative prompting and revision, so definitive attribution to human or machine sources is unreliable (Walton et al. 2025); isolating a stable "signal" of technological authorship struggles to distinguish unethical delegation from responsible engagement. Pedagogically, detection asks whether GenAI was used rather than how decisions were made, leaving invisible the evaluative work students already do when they reject inaccurate outputs, revise drafts and align content with disciplinary criteria. The ethical cost is a surveillance posture in which instructors "may be compelled to interpret probabilistic outputs without shared evidentiary standards, shifting their role toward investigation rather than assessment" (Lynch et al. 2021). Sharma does not argue for eliminating detection: it may serve as one layer of AI Governance for clear misrepresentation or deliberate outsourcing of intellectual labour, but cannot be the primary infrastructure of integrity.
  3. Assessment is where integrity becomes visible. Assessment is reconceptualised as an integrity apparatus: it signals what counts as evidence, how responsibility is interpreted and which academic behaviours are valued (Sadler 1989; Torrance 2007). Integrity-oriented assessment is not the same as Authentic Assessment, and the paper treats the difference as epistemic. Authentic assessment asks whether a task mirrors real-world practice; integrity-oriented design asks whether Learners can justify decisions and assume responsibility in relation to disciplinary standards. Integrity thus becomes an explicit, assessable criterion embedded in the task architecture rather than an incidental by-product of realism.
  4. Judgement, not reflection, is the evaluative core. Sharma separates judgement from reflective practice, which considers one's thinking retrospectively (Schön 1983), and from Metacognition, the awareness and regulation of cognitive processes (Zimmerman 2002). Judgement is evaluative (Sadler 1989): learners evaluate options against standards, justify choices within epistemic and ethical frameworks, and accept consequences under uncertainty, which places it at the intersection of cognition, ethics and professional accountability, much as professional judgement in regulated fields rests on reasoned decision-making rather than procedural compliance. Integrity is then not the absence of assistance but the accountable exercise of judgement in relation to shared academic norms.
  5. Integrity artefacts are evidence, not proof. Requiring learners to document reasoning risks privileging those more fluent in reflective discourse, which Sharma names as the danger of "the replacement of one compliance regime with another." Artefacts need contextual interpretation, shared criteria and professional calibration, and judgement as evidence of integrity "remains relational and situated rather than mechanically verifiable," so designing for visible judgement demands sustained educator reflexivity.
  6. Responsibility is redistributed, not discharged. Institutions articulate expectations that recognise GenAI-augmented authorship as normative; educators design assessments that make ethical decision-making visible; learners articulate judgement against disciplinary criteria. Integrity becomes a property of educational systems rather than of individual behaviour, aligning with international guidance advocating pedagogical rather than punitive responses to AI in education (Fengchun and Cukurova 2024).

Four practices that make judgement visible

These are practice-informed illustrations from the author's teaching context, selected because they represent the paper's theoretical commitments; the paper presents them as framework development, not as a validated intervention.

  1. Annotated decision trails. Learners document how GenAI contributed to their work and explain why particular outputs were retained, modified or rejected against disciplinary criteria, rather than cataloguing every instance of use. Integrity is evidenced through transparency and disclosure instead of inferred through detection, with attention falling on the thinking behind choices (Moya et al. 2024). Scalability in large enrolments is handled by requiring focused annotation of selected decision points, with rubrics attending to quality of reasoning rather than volume of reflection.
  2. Verification and accountability practices. Where GenAI contributes claims, examples or references, learners show how claims were checked, which sources were consulted, and how discrepancies or uncertainty were addressed. Inaccuracies become objects of responsibility rather than evidence of misconduct, a direct response to the probabilistic character of generative output. Sharma links these practices to designs that triangulate evidence across paired tasks to verify learning and strengthen assessment validity in AI-vulnerable contexts (Roe et al. 2025), and to formative assessment research on self-regulation (Nicol and Macfarlane-Dick 2006).
  3. Oral defence and dialogic accountability. Prompts ask where and how GenAI supported the work, which outputs were revised or rejected and why, and how the final product might be adapted to a specific context or constraint. Learners do not prove independence from technology; they demonstrate ownership of meaning, limitations and implications, and acknowledge uncertainty in both GenAI output and their own reasoning. Sharma argues that responsiveness to probing questions and grasp of disciplinary ideas are more robust indicators of learning than similarity metrics, and that criteria emphasising explanation rather than linguistic polish make oral defence a route to equity. Feasibility is met through brief structured conferences, small-group dialogic checks or selective sampling with shared calibration among instructors.
  4. Draft differences and version history. Learners compare selected versions to explain what changed, what role GenAI played at each stage, and how revisions improved clarity, accuracy or disciplinary alignment. Responsibility is shown not through a single "untainted" artefact but through how evaluative decisions evolved. Because learners choose which revisions to foreground, expectations for specificity and justification carry the weight that exhaustive tracking otherwise would, and grading attends to progression and coherence of final decisions rather than the quantity of revisions.

Policy, equity and trust

Institutional responses emphasising prohibition, detection and enforcement treat integrity as a problem to be managed through surveillance or policy compliance, which Sharma argues is increasingly misaligned with GenAI-augmented work. Detection tools operate probabilistically and are often misread as definitive evidence rather than indicators requiring contextual judgement, and regimes built on them can erode Trust, particularly when expectations for acceptable use remain ambiguous or inconsistently applied.

The equity argument is load-bearing. Learners engage generative tools from unequal positions shaped by language proficiency, prior opportunity, disciplinary familiarity and access to institutional support, and for some GenAI functions as linguistic or cognitive Scaffolding rather than substitution. Detection-centred models may therefore disproportionately affect learners who rely on generative tools for linguistic or cognitive support (James and Andrews 2024; Francis et al. 2025), and where integrity is inferred from surface textual features rather than articulated reasoning, inequities are more likely to be penalised than addressed. Research showing that norms of authorship and textual ownership are historically and culturally situated, and that plagiarism is interpreted differently across educational systems (Abasi and Graves 2008; Bretag 2016), supports the claim that integrity is a culturally situated practice requiring explicit dialogue rather than universal procedural rules, which is also where equity and integrity claims meet. Assessment designs that invite articulation of responsibility signal institutional confidence in learners, aligning integrity with learning and professional formation rather than suspicion.

As an illustration rather than data, Appendix A reproduces a decision trail from the paper's own production: the author asked a GenAI tool to role-play an editor and propose revisions, then recorded accepted suggestions (clarifying the contribution beyond Eaton 2023, defining judgement early, distinguishing the examples from authentic assessment, shortening the title) alongside rejected ones, among them a suggested paragraph contrasting Sadler, Schön and Zimmerman.

Limits of the argument

Sharma states plainly that the paper reports no empirical findings and claims no causal relationship between assessment design and learner behaviour, so the practices should be read as theoretical and illustrative. The reflective positioning in one supportive higher education setting introduces possible confirmation bias, and the designs may not transfer to differently resourced or less supportive environments. Interpretive reliability is the sharper internal problem: judgement is situated and its articulation varies across learners, so designs that require justification depend on educators' capacity to interpret reasoning consistently, raising questions of evaluative coherence across markers and disciplines. Practices may also become performative, with learners aligning discourse to expectations without internalising the commitments; making judgement visible creates the conditions for ethical dialogue without guaranteeing ethical practice. The empirical work the paper calls for is how learners respond to these designs, how reliably educators evaluate articulated judgement across disciplines, and how the approach fares under different resource conditions.

Connected Concepts

  • Academic Integrity — The paper's subject: integrity reframed as pedagogical practice rather than compliance
  • Evaluative Judgment — The construct the argument installs as the evaluative locus of integrity
  • Assessment — Positioned as the integrity apparatus through which judgement becomes visible
  • Assessment Validity — The design goal served by paired verification tasks
  • Authentic Assessment — The adjacent movement the paper distinguishes its designs from
  • AI Detection — Recalibrated from primary infrastructure to one layer of governance
  • AI Use and Disclosure Statements — What annotated decision trails ask learners to articulate
  • Generative AI — The technology whose outputs and errors become objects of verification responsibility
  • Metacognition — Distinguished from judgement in the paper's conceptual work
  • Self-Regulated Learning — Formative assessment and criteria transparency as the mechanism behind verification practices
  • Trust — What surveillance-oriented regimes erode
  • Equity — Detection-based inference penalises learners who use GenAI for linguistic or cognitive support
  • AI Governance — Institutional policy reoriented from enforcement toward assessment architecture
  • Higher Education — The sector the argument addresses

Connected Articles

Citation

Sharma, S. (2026). Educational integrity in GenAI-augmented assessment: making judgement visible. International Journal for Educational Integrity, 22(6).

Connected FAQs

Embed this page

Copy the code below to embed a chromeless version of this page in a learning management system or other website. The embedded view hides the site header, navigation, and footer.