On this page

Assessment validity — whether assessments measure what they claim to measure. AI in Education raises fundamental validity questions: do AI-graded assessments assess student learning or AI prompting skill? Does AI use invalidate traditional assessment assumptions?

Questions to Consider

  • Validity asks whether an assessment measures what it claims to measure. Before reading, if you saw a student submit a polished essay you suspected was AI-assisted, would you think the bigger problem was cheating, or that the task was no longer measuring what you thought it was measuring?
  • This page poses a sharp question: when students use AI, does the score reflect student knowledge or AI-prompting skill? Can you think of an assessment you've designed or taken where the score might now be telling you more about the tool than about the learner?
  • A key finding is that the same learner input can receive semantically different replies depending on which underlying LLM is used — introducing 'construct-irrelevant variance' that threatens reliability and fairness. If two students get different AI support purely because of the model behind it, how fair is the resulting comparison?
  • The page argues that even when an LLM scores well, transferring human score interpretations requires similarity in the latent structure of responses — and LLMs diverge from humans here. What does this suggest about trusting an AI that 'passes' an exam designed for humans?
  • Rather than trying to detect AI use, the knowledge base argues for redesigning assessments so they stay valid for AI-capable students. Why might redesigning the task be a more validity-preserving strategy than policing whether AI was used?
  • AI now serves as test-taker, test-maker, rater, and analyst — making the interpretive chain opaque. When every role in an assessment is filled by AI, what does it even mean to say an assessment is 'valid' for the human learner in the middle of it?

Introduction

Validity challenges

  • Construct validity: When students use AI on assessments, does the score reflect student knowledge or AI capability? Performance vs. learning research addresses this directly.

  • Cross-LLM construct-irrelevant variance in conversation-based assessment: Hao (2026) shows that even for a single conversational turn, the semantic content of Large Language Models (LLMs)-generated replies varies across models and conversational-context conditions. Within-model similarity consistently exceeds between-model similarity (0.715–0.795 vs. 0.443–0.604), and adding chat history meaningfully changes response content (median cross-history similarity ~0.40–0.45). Because the same learner input can receive semantically different replies depending on the underlying model, prompting and context alone cannot preserve response consistency — introducing potential construct-irrelevant variance that threatens validity, reliability, and fairness. Maintaining consistent assessment conditions as LLMs evolve is therefore an infrastructure challenge (symbolic rules, response templates, validation layers), not merely a Prompt Engineering one.

  • Evidence that is present yet never reaches the scorer: Abreu, Stari and Martí (2026) name a construct-irrelevant threat that precedes any reasoning: in AI review of experimental physics laboratory reports, an equation, graph, table or unit may be correctly included in a report yet not be retrieved from the processed content the model works from, and because a changed sign, value or unit alters the physical interpretation, a disagreement with the instructor need not indicate a reasoning error. Their remedy is procedural — high-resolution PDFs rather than photographed scans, equations written in an equation editor, legible axes and units inside figures, and instructions requiring a concrete citation for every score — because document quality sets the ceiling on whatever validity claim the resulting score supports.

  • Latent-structure validity across humans and LLMs: Strugatski et al. (2026) add a deeper validity condition: even when an LLM scores well, transferring human score interpretations requires similarity in the latent structure of responses. Comparing six Multimodal AI LLMs to human cohorts on chemistry and quantitative-reasoning instruments, they find LLM–human factor structures consistently diverge (LLM–human congruence below the human–human baseline), so performance on a human-normed exam is weak evidence about LLM abilities on the constructs the items were designed to measure.

  • Consequential validity: Do AI-mediated assessments have fair consequences? Language bias studies show that AI scoring can disadvantage non-native speakers.

  • Rule scope is itself a validity property: Wright (2026) argues that the prohibitions written across higher education since 2023 bar "generative AI" without the technical precision to separate the generation of assessed content from the conversion of a format, since optical character recognition, handwritten text recognition and speech-to-text are recognition technologies that infer what is already present rather than producing new content. Where a rule attaches to platform identity rather than function, it is over-inclusive by definitional accident: it captures transcription-only use of a multi-functional tool and sanctions a student who did not do the thing the rule was designed to prevent, falling hardest on Learners who rely on those tools for accessibility — an inference the assessment cannot support about the construct it claims to measure. He proposes function-based drafting plus four operational criteria (fidelity, non-augmentation, traceability, attestation) for borderline cases, and treats the reliance on detector output and style heuristics as evidence of how weak the rule's own evidential basis is.

  • Agentic completion removes the human-production assumption: Hadjisolomou & El-Haddad (2026) extend the validity analysis from generative assistance to agentic completion: autonomous AI agents can now log into an LMS, read materials, and complete unproctored asynchronous work end-to-end (demonstrated on a live course — a quiz scored 10/10 in under 5 minutes, and a fabricated-but-credible discussion-board reflection). Placing a "human-production assumption" at the base of Kane's argument-based inference chain, they show agent completion silently removes the backing for the scoring inference on which generalization, extrapolation, and decision inferences all rest — so every unproctored asynchronous score, including honestly earned ones, loses its interpretive support because authorship is unverifiable. Their decisive move is classifying this as a validity failure rather than an integrity one: an institution can punish misconduct and still lack grounds for the scores it reports. Detection is structurally insufficient (classifiers unreliable and biased; LMS monitoring sees the same clicks a student would), so the remedy is assessment redesign for verified human presence (a short oral component, process-visible drafts, class-session-specific references), with an equity-preserving menu of options rather than a proctoring mandate.

  • Construct-irrelevant variance in human grading of GenAI-assisted work: Luo & Dawson (2026) provide a direct empirical demonstration that human grading of GenAI-assisted work is shot through with construct-irrelevant variance. In scenario-based interviews with 33 university teachers, grading decisions were driven by person-oriented (student honesty, diligence), capability-oriented (independence from AI, GenAI skill, disciplinary mastery), relation-oriented (trust built with students), and justice-oriented (fairness, beneficence) values — all of which can vary grades on factors unrelated to the outcomes being assessed. Marking down GenAI-assisted work is justified, they argue, if and only if the AI use prevented students from demonstrating the assessed outcomes; otherwise value-driven grading threatens validity. The study grounds the validity framing in real teacher practice and calls for "two-way transparency" — teachers clarifying how GenAI use will affect grades, not just students declaring use.

  • Variation-at-scale demands construct-equivalent variant generation: VARIA (Lee 2026) subjects the "variation-at-scale" premise of AI-Integrated Authentic Assessment (AIAA) — replacing surveillance proctoring with per-student task variation — to an empirical, falsifiable check. Because each examinee receives a unique-but-equivalent performance task, the integrity guarantee is conditional on LLMs generating variants that are simultaneously surface-distinct, construct-equivalent, rubric-applicable, and difficulty-matched. Benchmarking three frontier model families across four prompting strategies on 600 variants, VARIA finds frontier generators satisfy the joint integrity criteria only at the margin (joint score 0.81–0.88) while non-frontier references collapse (0.50–0.55), and no single prompting strategy dominates all four properties — so "variation-at-scale cannot be solved by prompting alone" if the diversity threshold is set aggressively, and institutions must validate their specific model-prompt pair.

  • Under-grading bias in HITL AI scoring is a validity feature, not a bug: Curi et al. (2026) show that in a large-scale national writing assessment the systematic conservative (under-grading) bias of an LLM scorer, while problematic as a final decision, is precisely what enables safe delegation: AI-marked "passing" responses can be accepted with confidence while AI-marked "failing" responses (15.3–16.5% of cases) are routed to expert review — keeping the validity threat of AI error off the final outcome. Their IRT-and-Bookmark pipeline makes the proficiency-level impact of AI scoring legible rather than treating raw score agreement as the sole validity signal.

  • AI-driven grade inflation as a validity threat: Chirikov (2026) identifies a novel, technology-driven mechanism of grade inflation operating upstream of grading — on the production of graded work. In a difference-in-differences study of 500,000+ grades across 319 courses (2018–2025), courses with more AI-exposed tasks (writing, coding) saw the share of A grades rise by 13 percentage points after ChatGPT's release, with grade-distribution compression. A triple-differences analysis shows the effect concentrates in homework-heavy courses — evidence that AI task displacement (AI performing graded tasks before instructors observe them) inflates grades without a corresponding rise in skill. This reduces the comparability of grades across courses and erodes the informational value of transcripts in ways difficult to detect from grade distributions alone.

  • Program-level pass/fail susceptibility, and the marking criteria that set it: Ivory et al. (2026) judged ChatGPT output for every coursework assessment in a three-year psychology degree (40 assessments, 16 types) with a binary pass/fail, the decision markers actually make on a first read: 36 of 40 passed. Two validity implications follow. First, the operative variable is the pass boundary, not detection — marking that rewards structural fluency and the correct choice of analysis while condoning incorrect values lets fabricated statistics and hallucinated references clear a pass, so the assessment certifies tool output as student understanding because of how it is marked rather than because anyone was fooled. Second, multiple-choice items leaked their own answers: ChatGPT answered statistics questions correctly without the figure or output table shown to students, because the item stem and options implied the answer, which makes answer-option design a validity property of the item rather than a property of the AI.

  • Accuracy vs. agreement as distinct validity signals: Falahat, Das, Bhaumik & Thambi (2026) graded a 21-item pharmacy exam with ChatGPT-5 and found that moderate percent accuracy frequently coexisted with low concordance-correlation coefficients — a methodological distinction between scoring accuracy and agreement that limits AI's reliability as a grading substitute even where raw accuracy looks acceptable. Agreement was strong on objective items (CCC 0.935–1.000) but near-zero on short-answer and modest on essay (0.341–0.854), and rubric provision did not consistently close the gap.

  • Validity of AI-generated items: Assessing AI-Generated Exams shows that AI-generated questions, validated via Bayesian IRT, achieve difficulty and discrimination on par with expert-written standardized-exam items (reliability 0.79 vs. 0.72) — supporting the validity of course-tailored AI-generated assessments when backed by psychometric evaluation.

  • Authentic assessment: Authentic Assessment and the AI Assessment Scale propose validity-preserving assessment redesigns.

  • Confidence and calibration: Confidence-aware systems improve validity by flagging uncertain assessments.

  • Embodied and multimodal evidence: speech-only assessment can mistake verbal fluency for conceptual knowledge; Morphew et al. show that computer-vision gesture analysis coupled with LLM speech analysis increases construct validity and Equity by capturing understanding expressed through gesture, not just words — reducing bias against learners who express understanding non-verbally.

  • Homework stops certifying capability when a model can solve the task: Mikhasenko et al. (2026) document a concrete invalidation of homework as a capability measure in a Bochum particle-physics course: once a generative model can produce a correct solution, a submitted derivation no longer establishes that the student can solve the problem independently. Their response separates the two functions — research-shaped, AI-permitted homework kept as exploratory, bonus-bearing work, and a tools-free written examination made the sole determinant of the final grade — a validity-driven division of labor rather than a detection regime.

  • Decision-oriented validity for AI scoring at scale: Curi et al. (2026) illustrate a decision-oriented view of validity in the Acredita EB national writing assessment: rather than asking only whether AI scores agree with humans, they ask whether AI errors can change a certification decision. Automated IRT/Bookmark cut scores closely reproduced the operational cut scores, and systematic AI under-grading was neutralized by routing failing AI results to human review — notably against a reference standard that is itself contested, since ten expert raters scoring the same 50 texts never reached unanimity on any rubric item.

  • Student-side validity concerns about AI as grader: AlGhamdi (2026) adds the learner's voice to the validity question: 13 computing students whose handwritten writing task was scored by ChatGPT questioned whether an AI scorer can validly interpret intent, effort and institutional grading norms — one asking pointedly, "If ChatGPT [is] checking the exams, why are we going to university?"

  • Permission does not by itself degrade the task, and use rates are the wrong validity signal: Zou et al. (2026) surveyed 85 student teachers whose assessments explicitly permitted generative AI and found assessment engagement high and statistically unaffected by whether students used it (mean 4.21/5, SD 0.46; Mann–Whitney U = 781.5, r = 0.07, p = 0.536). With 62.4% (53) declining to use generative AI at all and adopters' use confined largely to proofreading (43.8%) and clarity checks (34.4%), the adoption rate carries little information about whether the assessment still warranted an inference about learning — and non-use driven by fear of wrongful plagiarism accusation (41.5% of non-adopters) is not evidence of validity either.

  • Sampling-frame validity when a school-level test is read as system evidence: Restrepo Morales et al. (2026) give the inference from scores to claims about a system a quantitative treatment. A pilot assessment of 1,198 volunteer students in 171 Salvadoran schools on the PISA-based Test for Schools produced results announced as comparable to Germany and Sweden, but the instrument produces estimates for a school and is not administered under the national sampling frame with its response-rate standards, exclusion limits and weighting — so a comparison between a school-level estimate and a country mean carries at least three unreported sources of uncertainty (school-level sampling error, country-mean sampling error, and the linking error of the scale equating). Rather than test the learning claim, the authors bound it: with the top 5.5% of the Salvadoran mathematics distribution matching the German mean under zero learning, and voluntary participation understood to bias achievement testing upward, the published evidence cannot separate a real effect from a selected sample. This is a failure of the comparison's validity inference, not of the students or of the instrument.

A conceptual proposal raises a validity boundary that applies to every AI-mediated assessment on this page: El Khoury and Ma (2026) argue that speed is not validity, and that AI-generated prompts, examples and transcript-based reports still have to be tested for accuracy, cultural responsiveness, Accessibility, interpretability and alignment with course outcomes. They also widen what counts as evidence: where dialogue becomes the assessed artifact, the transcript is a process record in which judgment, empathy, clarification and shared decision-making unfold in context, and the design question becomes whether that evidence supports the inference drawn from it — the same question automated scoring faces, answered by alignment among task, feedback and evidence of learning rather than by technological novelty.

  • AI-assigned item metadata is not psychometric evidence. Across a 10-week data science study of 311 deployed multiple-choice items, LLM difficulty ratings tracked the model's own Bloom labels (rho = 0.90) but not empirical item difficulty (rho = 0.06), so item difficulty has to be established from response data rather than from the generating model's labels (An & Wang, 2026).

Redesign over detection

The knowledge base argues that maintaining assessment validity requires redesigning assessments for AI-capable students, not detecting AI use. A parallel validity problem runs through research measurement: many AI-in-education claims rest on self-report data, which cannot support an inference about learning or competence however well the instrument itself is validated. Beyond detection approaches and Assessment represent validity-forward thinking.

Take-home products and the qualitative/quantitative split. Brunnström and Palmqvist (2026) give the redesign argument a specific proposal grounded in the SOLO taxonomy. Working through a cognitive-science take-home exam question with a chatbot, they find GenAI is strongest exactly where the taxonomy is lowest — producing comprehensive, fluent, factoid-type content at the quantitative (multistructural) level — while the qualitative level (relating, evaluating, generalizing) emerged only through repeated learner-driven calibration. Their recommendation is therefore differentiated by construct: take-home assessments should emphasize evidence of qualitative understanding, while quantitative recall-based knowledge is better assessed in class where GenAI is unavailable. The demonstration also underlines that a polished submitted artifact is weak evidence in either direction, and that valid redesign must account for how demanding legitimate AI-supported learning turns out to be (Summative Assessment, Academic Integrity).

A design-level counterpart comes from the same population the redesign argument usually ignores. Zou et al. (2026) found that reflective and personalized tasks were widely judged too personal for AI help, while a knowledge-heavy assessment in another course prompted strong intent to use generative AI for references — so what students did with AI tracked the construct the task asked them to evidence. The lesson for validity is precise rather than celebratory: design can remove the payoff from generic substitution, but the same study warns against reading an opt-out cohort as proof that it did, since a substantial share of non-use was attributed to fear of wrongful plagiarism accusation rather than to the task resisting AI.

Sharma (2026) supplies a redesign argument on the integrity side that turns integrity evidence into validity evidence. Extending Eaton's (2023) postplagiarism framing into assessment design, he argues that detection- and verification-based integrity models are misaligned with work in which human judgment and machine generation are entangled, and reframes integrity as a pedagogical practice enacted through Evaluative Judgment — the capacity to weigh options, justify academic choices and assume responsibility under epistemic uncertainty. The design consequence is that integrity should be evidenced rather than inferred: annotated decision trails, verification of GenAI-contributed claims, oral defense and version history are offered as integrity artifacts, and he is explicit that the difference from Authentic Assessment is epistemic — authenticity asks whether a task mirrors real-world practice, integrity-oriented design asks whether learners can justify decisions and assume responsibility against disciplinary standards, so integrity becomes an assessable criterion embedded in the task architecture. He names the validity risk of his own proposal rather than leaving it implicit: requiring documented reasoning privileges learners fluent in reflective discourse and risks "the replacement of one compliance regime with another", because judgment as evidence "remains relational and situated rather than mechanically verifiable" — the same interpretive-reliability problem this page records for every AI-mediated assessment.

Weidlich (2026) reframes redesign as a question of which inference a change is meant to protect. Working from Kane's argument-based validity, he separates five pressures that debate tends to collapse into one: construct underrepresentation, construct-irrelevant variance, attribution of performance, conditional extrapolation, and unsupported score use. AI use is not automatically a threat, because tool use is germane where the intended construct is AI-supported professional judgment and bypasses the target performance where it is unaided reasoning, so construct and AI conditions have to be specified together. His caution for redesign follows, since interventions trade one pressure for another: oral defenses may strengthen attribution while reducing reliability or accessibility.

Connections

Assessment validity connects to Authentic Assessment, Automated Grading, Confidence Aware AI Assessment, Formative Assessment, Academic Integrity, and RCT (which relies on valid outcome measures).

AI challenges validity at the epistemic level: Hathcoat, Slotnick & Miller (2026) argue that when LLMs serve as test-takers, test-makers, raters, and analysts, the interpretive chain becomes opaque and the object of measurement loses definition — reframing validity as requiring AI-fluent "cyborg" judgment, and Green et al. (2026) show AI scores can align with human raters (87% checklist) while the underlying rationale diverges, especially on measurement quality and weak reports.

When AI generates the assessable artifact. Kumar, Wongsirichot and Nanthaamornphong (2026) give the validity problem its sharpest disciplinary case: in computing education the AI produces the artifact being graded — the source code — so tool use, learning and assessment fuse into a single interaction, and a submitted codebase no longer separates learning from delegation. Across 72 studies they find efficiency gains that do not transfer to unaided performance (21 studies), Prior Knowledge moderating whether AI help becomes skill (6 studies), and detection research effectively absent from the evidence base (3 studies) while assessment redesign is comparatively well evidenced (25 studies). Their conclusion for practice is construct-specific and low-tech: add an oral component or other process-visible element to at least one high-stakes assessment per course — the single highest-leverage intervention in the review — and require critical engagement with AI output as a graded, observable component rather than an optional disposition (Academic Integrity, Assessment).

Validity under imperfect information

The sharpest recent reframing treats the generative AI problem as an evidentiary one. The student knows how a piece of work was produced; the institution observes the artifact and, at best, partial traces of the process — a product–process gap that Mohamed and Temimi (2026) formalize as assessment validity under imperfect information. On this account the question is not whether a rule was broken but whether the assessment still generates credible evidence of student reasoning, effort, and judgment, and each institutional mechanism — prohibition, monitoring, disclosure, redesign — is evaluated by which student response it makes most attractive. The validity lens also dissolves a false separation: integrity and validity are the same problem seen from different ends, because a finding of misconduct is itself a validity claim about what the work evidences.

A learner-side framework names the inference targets an artifact cannot settle by itself. Sak (2026) proposes attributional validity for judgments of gifted potential from AI-assisted work: the relation from capability to product is many-to-one, because learner competence, model capability, the learner's direction, human–AI fit and context combine differently to yield equivalent artifacts, so the product cannot identify the capacity behind it. The framework separates four targets of inference — independent competence, intellectual Learner Agency, hybrid capability, and developmental carryover — and names three errors that follow when one target's evidence is read as another's: overattribution, developmental illusion, and presumed capability equivalence. Its demand on the record mirrors the evidentiary discipline this section draws from the misconduct literature: name the capability a judgment is about, document how the work was produced (model and version, interface, permitted functions, adult mediation, access conditions), and require later evidence of transfer before reading assisted performance as a change in the learner.

Teichmann (2026) draws the procedural consequence: where prohibited use cannot be detected, a regime that still accuses on detector scores cannot warrant the inference it draws, and so produces unfairness without effectiveness. He argues for replacing the forensic question — did the student use AI? — with a validity question — did the student demonstrate the capability the task was designed to certify? — which relocates institutional effort into program-level assessment across linked tasks, oral and supervised elements at certification points, and Evaluative Judgment as an explicit object of assessment.

The same section of the evidence base supplies the case-file counterpart, and it is more uncomfortable than the detection literature alone. Munoz et al. (2026) coded every generative-AI misconduct case at one regional Australian university over three years — 1,162 cases carrying 1,855 discrete evidence items — and rated each item for probative value on relevance, credibility and inferential force. The strongest categories were the ones that do not depend on probabilistic classification of text (student admissions, observed prohibited exam behavior, and independently verified fabricated references), while detector output drew the weakest ratings of any category: similarity or Turnitin reports were wholly weak in inferential force and standalone detector outputs wholly low in credibility, falling from 8.8% of items in 2024 to 0.5% in 2025. The validity failure is procedural as much as instrumental — no stage of the pipeline set a minimum evidentiary threshold or required investigators to weigh probative value before progressing an allegation, so evidence quality showed no reliable relationship to case outcomes, which is exactly the evidential standard issue catalogued on Legal Issues and Risks. Their proposal to apply a credentials framework at the point of allegation rather than only at determination is a validity demand: an allegation is a claim about what the work evidences, and it should be warranted by materials capable of supporting it. Hadra, Cambridge and Mesbah (2026) show how far the instrument is from that standard, with macro accuracy of 0.69 (Originality) against 0.61 (Turnitin) across 192 texts, near-total failure on hybrid human–AI writing (sensitivity 0.02 and 0.31), declining accuracy with text length and on scientific writing, and a borderline tendency to misclassify EFL student writing as AI — so a detector score cannot establish the fact a misconduct finding asserts.

The human-marking benchmark and its ceiling

The OpRaise comparison of three frontier models against 761 authentic essays across three UK universities makes the benchmark part of the validity argument. Human marks were used as ground truth on the explicit ground that academic judgment is the socially accepted standard, while the authors acknowledge that human markers agree only moderately with one another — which caps the AI–human agreement that could reasonably be demanded, and means a correlation cannot be read as ready-or-not without a reference point. Within that frame the failures appeared as systematic structure rather than random error: marks compressed toward the middle of the scale, so the best and worst essays were misjudged most; agreement was weakest at grade boundaries; and AI marks tracked vocabulary range, connectives and sentence complexity while human marks were broadly insensitive to them. The practical lesson is double-edged, because the same study found reliability to be excellent — identical re-marks across time and high agreement between models. A validity case for automated marking therefore cannot rest on stability or on average agreement; it has to show the absence of systematic deviation, which is exactly what this evidence did not find.

  • Agreement and benchmark error masquerade as learning. Two 2026 studies show distinct routes by which a score can look valid while establishing something else. An open-ended marketing-writing study found LLM-human absolute agreement of only ICC(2,1) .435 and a hybrid that was significantly worse than the LLM alone, with anchor composition moving agreement from .338 to .902 (Agreement and error in automated scoring of student marketing posts). An expert audit of six physics benchmarks attributed 95.20% of audited rejections to defective items or graders rather than model error, moving CritPt mean@5 from 32.29% to 87.50% (How Good Are Frontier Models at Physics? Expert Re-Grading Reveals Broken Evaluations and Near-Saturation of Leading Benchmarks). In both cases the threat is construct-irrelevant variance outside the model being scored.
  • Agreement among LLM coders is consistency, not validity. On an expert-labeled educational dialogue corpus, a cross-model agreement filter retained only 33 of 74 corpus-derived behavioral assertions and 11 of 48 construct-derived ones, and shared model errors survived the filter, so inter-model agreement cannot stand in for construct validity (Bernado et al., 2026).

Connected Concepts

Connected Articles

Connected FAQs

Embed this page

Copy the code below to embed a chromeless version of this page in a learning management system or other website. The embedded view hides the site header, navigation, and footer.