π Research Article
Assessment Twins: An Approach for Strengthening Assessment Validity in the Age of Generative AI
Synthesis: Roe, Perkins & Giray (2026) introduce assessment twins as a practical approach to redesigning assessment tasks for the age of generative AI. An assessment twin pairs a GenAI-vulnerable task (e.g., a take-home essay) with a second, less vulnerable task assessing the same learning outcomes, scheduled closely to allow cross-verification β enhancing Assessment Validity without abandoning pedagogically valuable assessment formats. The paper maps GenAI threats across Messick's six strands of validity evidence and proposes a three-step design process (identify vulnerabilities, align outcomes/select the twin, develop interdependent marking). It directly addresses the Academic Integrity problem of GenAI, complementing detection-focused responses with a validity-driven, pedagogy-first design strategy for AI-mediated assessment.
Key Findings
- Assessment twins as a validity-focused response to GenAI. Two deliberately linked, interdependent components that address the same intended learning outcomes through different modes of evidence, scheduled to allow cross-checking against a known vulnerability (e.g., GenAI completion). The twin approach triangulates evidence across pedagogically valuable but GenAI-vulnerable formats.
- A systematic validity mapping. Using Messick's unified validity framework (via Shaw & Crisp's six strands β content, substantive, structural, generalisability, external, consequential), the paper shows how GenAI threatens each strand and how twinning mitigates those threats.
- Retaining formative value. Many GenAI-vulnerable tasks (take-home essays, research reports) carry substantial formative value. Rather than discarding them, twins retain the vulnerable task for learning while pairing it with a twin that supplies reliable summative evidence β the original task supports learning, the twin confirms achievement.
- A three-step design process. (1) Identify GenAI vulnerabilities (requiring assessor AI literacy); (2) consider learning outcomes and choose a complementary twin assessment; (3) develop an interdependent marking framework (confirmatory threshold or confirmatory weighting).
- Context-dependent application. Twins suit institutions that can support resource-intensive confirmatory tasks; in resource-limited, very large-cohort settings, a full redesign using established frameworks (e.g., the AIAS) may be more effective. Challenges include resource intensity, equity concerns, and the need for empirical validation.
The validity problem GenAI creates
GenAI's advanced capability to produce extended works leads to GenAI-assisted plagiarism ('Aigiarism'), making it challenging to determine whether students' work is their own and compromising assessment validity β since it becomes impossible to identify whether students have met course standards. AI detection was rapidly promoted as a remedy, but detection is now "all but impossible" and surveillance-focused responses can harm the relational dimension of assessment (Trust between students and institutions). Structural assessment redesign is a more robust and pragmatic response, though the literature offers few clear methods β a gap this paper addresses.
The assessment twins concept
An assessment twin comprises two deliberately designed, interdependent components that (a) address the same intended learning outcomes, (b) require different modes of evidence or production, and (c) are scheduled so performance on each can be cross-checked to mitigate a known vulnerability, thereby enhancing construct validity compared to either component alone.
The approach is distinct from traditional protocols like the oral viva voce in its core organising logic: it is fundamentally a validity-driven response to a GenAI vulnerability from the outset. The distinction between formative and summative assessment is central β GenAI-vulnerable tasks are retained for their formative value (the process of completing them is itself meaningful learning) while the twin provides reliable summative evidence of the same outcomes.
The design process in practice
- Step 1 β Identify vulnerabilities: requires assessor AI Literacy about what GenAI can/cannot do; remote unsupervised assessments are more vulnerable; tools like the PANDORA rubric help.
- Step 2 β Align outcomes & choose the twin: both components map to the same learning outcomes; complementary modes include real-time demonstrations, oral explanations, group discussions, or Q&A. Crucially, oral/video submissions are only as valid as the work underpinning them β live, interactive components (unscripted questions, real-time discussion, in-situ problem solving) mitigate pre-generated AI reliance.
- Step 3 β Develop interdependent marking: either a confirmatory threshold (minimum twin performance required for the original mark to stand) or a confirmatory weighting (e.g., essay mark Γ oral confirmation score / 2). The worked example pairs a take-home research essay with a 15-minute oral interview.
Context and limitations
Twins are most appropriate when a task is pedagogically rich but AI-susceptible. The approach is resource-intensive, requiring faculty time, administrative coordination, and institutional support. Cohort scaling matters: small groups (5β25) suit individual orals/vivas; medium (25β75) suit peer-group presentations and poster events; large (75+) may use random sampling with transparent, non-punitive selection or multiple-assessor peer discussions. In resource-limited, very large contexts, a full redesign using an established framework (e.g., the AIAS) is more effective. The framework requires empirical validation.
Implications for AI in education
Assessment twins offer a practical, validity-driven complement to the wiki's Assessment Validity, Academic Integrity, and Authentic Assessment threads β moving beyond detection toward structural assessment design that prioritises pedagogy while supporting meaningful learning outcomes. The approach connects to assessment policy choices about summative format and to the broader theory-building strand on how institutions redesign assessment for AI-mediated education, alongside cognitive stewardship and authentic assessment redesign.
Connected Concepts
- Assessment Validity
- Academic Integrity
- Authentic Assessment
- Formative Assessment
- Summative Assessment
- AI Literacy
- Generative AI
- Higher Ed
- Plagiarism Detection
- Educational Policy AI
- Theory Development AIED
Connected Articles
- Credential Cognitive Stewardship AI Assessment β Cognitive stewardship for AI-mediated assessment
- AI Assessment Scale Reform β AI assessment scale reform
- GenAI Assessment Governance β GenAI assessment governance
- Beyond Detection Authentic Assessment AI 2025 β Authentic assessment redesign
- GenAI Declaration Frameworks Higher Education β GenAI declaration frameworks
Citation
Roe, J., Perkins, M., & Giray, L. (2026). Assessment twins: An approach for strengthening assessment validity in the age of generative AI. Journal of Applied Learning & Teaching, 9(2).