Research Article
Which inference is at risk? Assessment validity reasoning and generative AI
Synthesis: Weidlich argues that generative AI has not created one assessment problem but several that debate collapses together, and that treating them as academic integrity misses what is actually at stake: the meaning of the evidence. Working from Kane's argument-based validity, he sets out a case where a student submits a strong take-home essay and then performs poorly in a supervised oral explanation, and shows that the discrepancy alone cannot tell you what capability the essay now evidences, whether the score reflects tool access or marker response, how much of the work belongs to the student, or what decision the combined evidence can justify. From there he separates five pressures on the inference chain, offers a five-step redesign sequence that starts with specifying the intended claim and the AI conditions, and gives a Validity Reasoning Matrix that links each concern to a threatened inference, a design question and the local evidence needed.
Key Findings
- The same symptom admits several diagnoses. A gap between an AI-permitted essay and a supervised oral establishes neither misconduct nor competence: it can indicate construct underrepresentation, construct-irrelevant variance, an attribution problem, a mismatch between performances, or an unclear rule for combining evidence.
- AI use is not automatically a threat to validity. If the construct is AI-supported professional judgment, tool use is germane to it; if the construct is unaided disciplinary reasoning, the same use bypasses the target performance, so construct and AI conditions must be specified together.
- Detection adds construct-irrelevant variance rather than removing it. Detector performance varies across tasks, disciplines and model versions, and disclosed or suspected AI use can act as a biasing cue for markers, so a security response can contaminate score meaning and fairness.
- Authorship is not the same as ownership of meaning. Because AI assistance is iterative and interwoven with drafting, who produced the words explains less than what claim the performance warrants, shifting the task from policing provenance to gathering complementary evidence of understanding.
- Valid under one set of conditions does not mean valid under another. An AI-supported coursework score may not support claims about independent competence, and an AI-restricted exam may underrepresent the capability that responsible AI use in professional settings requires.
- A defensible interpretation does not justify every decision. Even a bounded, well-supported claim may be insufficient for progression, certification or selection, which requires separate warrants for thresholds, weighting and the resolution of inconsistent results.
- Redesigns trade one validity pressure for another. Oral defenses can improve attribution while reducing reliability or accessibility, and authentic tasks can improve extrapolation while complicating scoring, so responses should be chosen for the inference they strengthen and checked for what they weaken elsewhere.
The redesign sequence and the matrix
The practical contribution is a five-step sequence: specify what a high score means, specify the AI conditions under which that interpretation should hold, locate the vulnerable inference by asking which warrant the concern calls into question, choose a design response while checking what it changes elsewhere in the chain, and identify what local evidence the stakes require. Proportionality governs the last step, with low-stakes formative tasks needing condition statements, exemplars and modest monitoring, while progression, award and certification decisions need secure complementary tasks, moderation, an explicit mapping between outcomes and the assessment mix, and a review of equity effects. The same matrix works at both levels: an instructor uses it to decide whether a task needs redesign, extra attribution evidence or clearer conditions, while a program can see whether a curriculum leans too heavily on polished take-home products. Applied to one case-analysis essay, it yields different redesigns for independent reasoning, AI-supported analysis, writing fluency, capability in target settings and formative use.
Where authenticity fits
Weidlich also repositions authenticity, which is often treated as either a competing principle to validity or a property that protects a task from AI. In this account authenticity is backing for particular inferences, especially domain representation and extrapolation: an assessment is authentic to the extent that it samples relevant features of the target practice under defensible conditions in a way that supports the intended interpretation and use of the score. That makes authenticity a claim needing warrant rather than a label that settles the argument, and explains why tasks can be authentic in scenario yet weak in attribution evidence, or faithful to current practice yet misaligned with educational aims. Consequences and incentives are handled as a cross-cutting design consideration rather than a separate inference, because a regime that rewards polished products over defensible reasoning has an incentive problem before it has a misconduct problem, and because measures introduced for security can change anxiety, accessibility, workload and trust.
What this means for practice
- Instructors. Write the claim and the AI conditions before choosing a task, and check each redesign for the inference it weakens rather than assuming that restriction solves the problem.
- Assessment designers and professionals. Use complementary evidence of the student's own control over AI-assisted work, such as justified recommendations, source-use explanations or a focused follow-up, instead of relying on provenance or detection.
- Administrators. Match the strength of local validation to the stakes, and treat equity effects of new controls as part of justification rather than an afterthought.
- Program teams. Audit the assessment mix for over-reliance on one kind of evidence, and ground high-stakes decisions in multiple sources rather than a single data point.
Limitations
- The argument is conceptual and design-oriented: the author states that the proposed validity reasoning approach is not a validated protocol and its practical value should be examined empirically.
- Whether educators and program teams can apply the matrix consistently, make better redesign decisions with it, and gather the evidence it asks for under ordinary workload constraints is untested.
- The framework can structure disciplinary deliberation but cannot determine whether AI-supported performance should be central, peripheral or restricted in a given field, and AI condition statements may be misunderstood or experienced as shifting responsibility onto students.
Connected Concepts
- Assessment Validity
- Academic Integrity
- AI Detection
- Critical Thinking
- AI Literacy
- Equity
- AI Governance
- Generative AI
- Large Language Models (LLMs)
- Higher Education
Connected Articles
- Coauthorship integrity: Reconceptualising assessment validity for the age of generative artificial intelligence — coauthorship integrity as a reconceptualization of assessment validity
- Assessment twins: An approach for strengthening assessment validity in the age of generative AI — assessment twins as a way to strengthen validity in the generative AI era
- Beyond Detection: Redesigning Authentic Assessment in an AI-Mediated World — redesigning authentic assessment rather than policing detection
- Assessment Design Under Imperfect Information: Generative AI, Disclosure, and Student Response in Higher Education — assessment design under imperfect information about AI use
Citation
Weidlich, J. (2026). Which inference is at risk? Assessment validity reasoning and generative AI. Assessment & Evaluation in Higher Education. Advance online publication.