Research Article
GenAI assessment and language equity: Drawing the line between support and substitution
Synthesis: Li argues that generative AI has collapsed the practical boundary between language assistance and substantive authorship, and that for students who use English as an additional language (EAL) the collapse is not neutral: integrity rules treating GenAI as a single category of unauthorised assistance convert linguistic disadvantage into integrity risk. The remedy proposed is a purpose-based, rather than tool-based, support-substitution boundary anchored in the assessment construct, developed through indirect discrimination reasoning, procedural fairness principles, and an argument-based account of validity. The article supplies a model policy architecture with calibrated disclosure, a three-stage operationalisation of proportionality, and assessment design levers that make reasoning visible without treating polished English as proof of authorship.
Core Argument
- The risk is structural, not behavioural. Li states explicitly that the argument is not that EAL students are more likely to behave dishonestly, nor that they should be allowed to outsource substantive work. It is that rules which do not distinguish language support from substitution impose higher compliance burdens on the students most likely to use legitimate support, raise suspicion where surface fluency shifts, and enable selective enforcement grounded in weak evidence.
- The boundary is defined by function and construct, not by tool. Tool lists go obsolete and encourage compliance by interface rather than purpose. Permitted language support covers surface-level interventions that improve readability without adding ideas, sources, or analytical structure: grammar, punctuation and sentence-level clarity edits, translation for comprehension, and first-pass drafting that the student substantively rewrites. Substitution covers interventions that create or materially reshape the intellectual work being evaluated: generating arguments or counter-arguments, applying disciplinary rules to facts, restructuring the analytical sequence, and generating citations for inclusion.
- Undifferentiated rules produce cohort-skewed burdens. An outright ban removes a scalable form of language support; a permit-but-disclose regime loads compliance work onto EAL students whose use is more frequent and iterative; and prohibiting editing beyond minor changes concentrates suspicion on writers whose fluency has shifted most. Allegations also carry reputational, academic, and sometimes visa or financial consequences.
- Detection-led reasoning cannot carry the evidentiary weight placed on it. Detector outputs may have a triage role but rarely constitute proof. Li's preferred posture is triangulation: staged submissions, draft histories, verifiable source trails, and a brief construct-aligned conversation.
- Proportionality is operationalised in three stages. At the threshold stage it decides whether a matter becomes a minor compliance issue, an educative intervention, or a formal allegation. At the evidentiary stage it governs the burden placed on the student to account for their writing process, with heavier obligations on the institution as stakes rise. At the sanction stage it maps outcomes to conduct, from warnings for non-disclosure of permitted support to escalating sanctions for deliberate substitution.
- The framework accommodates permissive regimes, under three conditions. Scaled models such as the AI Assessment Scale are compatible provided the use is expressly authorised for the specific task, the construct is re-specified so that what is assessed is clear to students and markers, and disclosure is calibrated to be feasible. What the framework rules out is undifferentiated permission, not breadth of permission.
Why the boundary has to be tied to the construct
Construct statements in higher education are often implicit: staff say they assess analysis or judgement, yet allocate marks for fluency, idiomatic expression, and rhetorical confidence. For EAL students this creates construct-irrelevant variance, where the score is partly driven by language form rather than the reasoning the task purports to measure. Li anchors the point in Kane's (2013) interpretation/use argument, which requires an explicit chain of inferences from performance to scoring, generalisation, and institutional decision, and in Messick's (1989) insistence that foreseeable consequences for different groups sit inside the validity question.
The line is illustrated with an essay in which a student has independently developed the claim that a housing policy produces unequal effects, chosen the evidence, and organised the analysis around three causal mechanisms. Asking a tool to correct grammar, shorten overlong sentences, or clarify transitions while preserving claims, evidence, paragraph order, and analytical relationships is language support, because the substantive architecture remains the student's. Asking it to decide which mechanism should lead, to combine two mechanisms into a new thesis, or to supply a counterargument is substitution, because form and substance are reshaped together. A rubric can separate the origin and quality of reasoning from the effectiveness of its communication, while recognising that major structural revision is substantive rather than linguistic.
Li also draws on formative assessment and Feedback scholarship (Sadler, 1989; Boud & Molloy, 2013) to argue that integrity is not protected only at the point of sanction, and that designs requiring students to evidence judgement over time serve the same end better than surveillance.
Equality logic, fair process, and reviewability
The article applies the reasoning structure of indirect discrimination analysis rather than surveying discrimination law: identify where a facially neutral rule predictably imposes cohort-skewed burdens, then ask whether those effects are justified by legitimate aims pursued through proportionate and practically workable means. Integrity and assurance of learning are legitimate aims; the question is necessity and proportionality once GenAI is present. Blanket prohibitions may be defensible where independent writing under specified conditions is part of the construct, but even then they depend on a clear rationale, accessible alternatives for language support, and designs that do not push staff toward speculative inference.
Fairness begins with notice. Generic statements that students must submit their own work give no workable boundary where the same interface delivers both permitted editing and prohibited drafting, so defensibility depends on ex ante specificity about permitted functions, prohibited functions, and required disclosure. Consistency and recorded reasons follow, supported by procedural fairness scholarship on the handling of suspected misconduct (Evans & Levine, 2016) and by findings that perceived fairness and respectful treatment shape whether students regard university authorities as legitimate (Główczewski & Burdziej, 2023). Li draws on the administrative-law requirement that decision-makers justify conclusions by reference to the relevant rule and evidence (Aronson et al., 2016), arguing that reviewability improves when an institution can state what rule was applied, what evidence was relied on, and why the outcome was proportionate, which is also a condition of legitimacy. This is the point at which an assessment rule stops being a Pedagogies and Teaching Strategies question: reviewability and procedural fairness are what an institution is judged on when a student challenges a finding, and Li's argument is that a rule ignoring language background fails that test before it produces a wrongful accusation.
Universities routinely permit human language editing, peer feedback, and learning support services. The accommodation analogy sharpens the Ethics: treating support as a site of potential cheating stigmatises the person rather than the conduct, and makes exhaustive disclosure of routine translation and editing use disproportionate.
Policy architecture, disclosure and implementation
The framework is delivered as three instruments. Box 1 gives working definitions and quick tests: permitted support requires all three of no new ideas, no new sources, and no material re-ordering of analysis; substitution follows if any of generating arguments or counter-arguments, applying rules to facts, restructuring the analytical sequence, or generating citations applies. Its edge-case column covers paragraph-level re-composition, translation prompts that request argument templates, and citation helpers that invent sources. Because the tests are shared, two markers should characterise the same conduct alike, narrowing the discretionary range even without perfect inter-rater reliability.
Box 2 provides calibrated disclosure templates: nothing for embedded low-risk functions such as basic spellcheck, a one-line statement for external language editing confirming that no new ideas, sources, or restructuring were added, and a process note of roughly 80 to 120 words where an assessment expressly permits limited ideation or drafting assistance. The one-line statement is deliberately shorter than a typical bibliographic entry, so EAL students using routine support do not face compliance costs exceeding those borne by monolingual peers. Box 3 is a decision rubric listing acceptable evidence types, weak proxies that should not be determinative (polished language or sudden fluency, non-native phrasing, single detection scores, and assumptions about what an EAL student should sound like), the balance-of-probabilities decision question, and sanction mapping.
Implementation rests on design levers rather than on detection: staged submissions across proposal, outline, annotated bibliography, draft and final; short standardised oral or in-class verification components that test reasoning without becoming a language performance test; and critique-based questions such as evaluating an AI-generated answer and correcting its authorities (Bretag et al., 2019; TEQSA, 2023). Li recommends four short checks in course approval: construct clarity, stated in one or two sentences; identification of construct-irrelevant barriers the institution will remove through permitted support; learning traces that make independent reasoning observable; and proportionality, since high-stakes high-discretion tasks concentrate fairness risk. A small random sample of students, plus cases showing material discontinuities between stages, can be invited to a short structured conversation about two or three substantive choices, with clarification or first-language support permitted; inconsistency triggers inquiry rather than an automatic misconduct finding.
Scope, limits, and the research agenda
Li is explicit that the analysis is not empirical: it does not measure student behaviour or policy comprehension, and makes no claims about rates of GenAI use by EAL students. It uses Australian regulator and sector guidance, principally from the Tertiary Education Quality and Standards Agency (TEQSA, 2023, 2024, 2025) and Universities Australia (2017), as illustrative reference points, and adopts a doctrinal and normative method rather than a jurisdictional survey. The framework is offered as testable rather than final, with three tractable avenues: comparative policy work on whether a purpose-based support-substitution boundary is associated with different referral rates, sanction patterns, and student comprehension across EAL and non-EAL cohorts; marker-calibration studies measuring inter-rater reliability when decision-makers apply the Box 1 quick tests and Box 3 rubric; and validity-argument case studies documenting whether institutional reasoning from student work to integrity finding survives appeal and review. The manuscript's own declaration records that ChatGPT and Copilot were used for language editing and to test alternative outlines, with the author retaining responsibility for arguments and sources.
Connected Concepts
- Academic Integrity — Fairness and responsibility are the values that anchor the language equity argument
- Equity — Language equity is reframed as a legality and rule-design problem
- Multilingual Learning — EAL cohorts are the population whose burdens the framework is built to reduce
- Assessment Validity — Kane's interpretation/use argument and Messick's consequential validity structure the boundary
- AI Use and Disclosure Statements — Disclosure is calibrated rather than exhaustive to avoid cohort-skewed compliance costs
- AI Detection — Detector output is demoted to a triage signal that cannot carry a finding
- Generative AI — The same interface performs permitted editing and prohibited drafting
- Assessment — Construct clarity and rubric design make the boundary workable
- Language Learning — Translation for comprehension and iterative editing count as legitimate support
- Educational AI Policy — The deliverable is a model policy architecture with clauses and decision rubrics
- AI Governance — Cohort-level monitoring closes the loop between integrity enforcement and quality assurance
- Culturally Relevant Pedagogy — Cultural specificity of authorship proxies is a reason to distrust surface signals
- Authentic Assessment — Staged tasks and construct-aligned verification generate evidence without surveillance
- Trust — Legitimacy depends on rules students can understand, comply with, and see as fair
- Legal Issues and Risks — reviewability, procedural fairness and the exposure created by imprecise rules
Connected Articles
- Reimagining Success and Failure: Equitable Assessment Practices in an Age of Artificial Intelligence — Equity as an explicit criterion for GenAI-era assessment design
- Generative AI and linguistic diversity in academic writing and publishing: Perspectives from World Englishes — Linguistic diversity and bias in AI-mediated academic writing
- Purpose Before Policy: Academic Integrity, Generative AI, and Rhetorical Stance — The parallel argument that purpose should precede policy
- Generative AI as a Design Variable: An Evidence-Centered Framework for Principled Governance in STEM Assessment — Institutional governance of GenAI assessment beyond single-tool rules
- Coauthorship integrity: Reconceptualising assessment validity for the age of generative artificial intelligence — Reconceiving assessment validity when authorship is distributed
- Assessment Design Under Imperfect Information: Generative AI, Disclosure, and Student Response in Higher Education — Disclosure regimes and what assessment can infer from them
- Marked Pedagogies: Examining Linguistic Biases in Personalized Automated Writing Feedback — Linguistic bias in writing feedback and marking
- The Integrity of Psychology Assessments in the AI Age: A Critical Examination — Programme-level evidence that the marking boundary, not detection, is the weak point
Citation
Li, G. (2026). GenAI assessment and language equity: Drawing the line between support and substitution. Journal of Academic Ethics, 24, 83.
Connected FAQs
- ❓ How Do I Redesign Assessment So That a Grade Still Tells Me Something Defensible About What the Student Knows or Can Do?
- ❓ How Do I Write a Course AI Policy and Communicate It to Students?
- ❓ How Should AI in Education Research Incorporate Equity, Accessibility, Privacy, Ethics, and Pedagogical Safety?