Research Article
Assessing Learning When AI Can Produce the Product: An Integrated Framework for Cognitive Ownership
Synthesis: Rothwell (2026) offers a conceptual framework for assessing learning when Generative AI can supply the reasoning behind a student's product. The paper distinguishes four things that are easily conflated: product quality, provenance, demonstrated capability, and learning as change. It defines cognitive ownership as a learner's demonstrated capacity to understand, evaluate, adapt, and accept responsibility for consequential reasoning under stated conditions of assistance, organized into eight indicators across four domains — understanding, epistemic judgment, adaptive performance, and reflective accountability. To make those judgments examinable, it proposes mutual evaluation: independent student and instructor evaluations of the same work against shared criteria, evidence-based comparison, and documented responses to disagreement, with agreement treated as neither necessary nor sufficient proof of learning. The design rests on argument-based validity and on evaluative judgment, and it preserves Self-Assessment as a formative resource while separating it from summative inference. A worked graduate leadership example shows task design, scoring weights, and how conflicting evidence should narrow or support a capability claim. Five propositions and a validation agenda address predictive value, task sampling, Feedback, fairness, and feasibility. The framework is proposed, not empirically validated.
Key Findings
- Product quality is not capability. A polished AI-assisted artifact establishes the quality of the work but not how much of its reasoning the learner understands, so Assessment must specify the capability it can legitimately claim.
- Cognitive ownership has four domains. Eight indicators — explain and reconstruct; verify, challenge, and judge; transfer and defend; reflect and account — are grouped into understanding, epistemic judgment, adaptive performance, and reflective accountability.
- Mutual evaluation structures judgment. Students and instructors evaluate the same work independently against shared criteria, then compare reasons; agreement is neither necessary nor sufficient evidence of ownership.
- Conflicting evidence should narrow claims. A strong product with weak verification supports a claim about the artifact while leaving individual capability uncertain, prompting targeted reassessment rather than an inference about ability or integrity.
- Assistance conditions must be stated. The framework asks assessors to identify where AI could substitute for the target capability and to increase evidentiary demand with the stakes and uncertainty of the decision, not with AI use alone.
- Fairness and feasibility shape design. Documentation should be selective, response modes must preserve the construct, and workload must be planned — an eight-minute verification for 120 students takes 16 contact hours before scoring.
- The framework awaits validation. Five propositions cover predictive value, evaluative judgment, verification value, evidence diversity, and accessible response modes, but no validated instrument or demonstrated effect is reported.
From product quality to cognitive ownership
The paper opens with a familiar problem: an excellent student paper can demonstrate the quality of a completed analysis without showing how much of that analysis the student understands, and generative AI widens the range of consequential reasoning that can be supplied externally. Rothwell treats this as a conditional problem rather than a universal threat, asking which capabilities a particular assessment permits an institution to claim, under which conditions of assistance, and with what evidentiary support. Drawing on argument-based validity and Distributed Cognition, he argues that coordination, verification, and decisions about delegation may themselves be legitimate capabilities, so assessors must name the unit of analysis — the individual, the human–AI arrangement, or both. Four distinctions organize the argument: product quality, provenance, demonstrated capability, and learning as change. An AI-use declaration, he notes, addresses only provenance and does not establish the other three.
Mutual evaluation and conflicting evidence
Mutual evaluation is the framework's central mechanism. Learners record an evaluation before receiving instructor comments; the instructor evaluates independently before seeing the student's ratings; comparison then focuses on consequential differences, and the record states what changed and why. The process has two functions that must be separated: formatively, comparison can improve the work, while summatively the quality of the learner's explanation is itself evidence of evaluative judgment. Rothwell warns of predictable vulnerabilities — students may imitate rubric language, generate a self-critique with AI, or agree to avoid conflict, and instructors may reward students whose style resembles their own. He also stresses that insufficient ownership evidence is not automatically evidence of misconduct, so integrity concerns and capability verification should follow separate decision processes. His worked graduate leadership example shows how a student's rating of a sophisticated readiness matrix can be probed against the case evidence, turning disagreement into an assessment of whether the student understands the matrix's evidentiary limits.
Fairness, feasibility, and the validation agenda
The framework insists that verification must not become an irrelevant barrier. Oral fluency, typing speed, and institutional vocabulary can distort assessment when they are not intended outcomes, so alternative response modes should preserve the construct and offer genuine Accessibility. Access matters too: a task rewarding capabilities available only through a paid model may partly assess purchasing power, and institutions can supply a common tool or permit equivalent non-AI pathways, an approach that advances equity. Documentation should be selective and proportionate, since full conversation histories may include personal disclosures, raising Privacy questions about who can access records and how long they are retained. Workload requires concrete planning — Rothwell calculates that an eight-minute verification for a hypothetical class of 120 students requires 16 contact hours before preparation, scheduling, and scoring. Sampling only students suspected of AI use, he adds, can make verification appear punitive, so selection should follow a disclosed neutral rule.
What this means for practice
- Assessment designers. Start from the claim you intend to support — independent interpretation, evaluation of AI recommendations, or accountable collaboration — and state the permitted assistance conditions, because these require different evidence.
- Assessment designers. Score product quality separately from demonstrated ownership; the paper's illustrative weighting is 40% product, 30% verification and adaptive performance, 20% mutual evaluation, and 10% reflective accountability, with a separate minimum where a capability is indispensable.
- Instructors. Run mutual evaluation in the right order — student self-evaluation first, instructor evaluation independently, then evidence-based comparison — and treat agreement as neither necessary nor sufficient evidence of learning.
- Instructors. Use outcome-aligned transfer and verification tasks, ideally with generative assistance excluded for that component, and follow a common question blueprint so difficulty does not depend on assessor improvisation.
- Faculty developers. Focus training on outcome definition, evidence selection, task variation, and interpreting disagreement; training limited to prompting or detection does not resolve what a grade represents.
Limitations
- The framework is a conceptual proposal with an illustrative example; no validated instrument, established threshold, or demonstrated treatment effect is presented.
- The author states that the selected literature grounds the reasoning but does not establish that all eight indicators are necessary for every task.
- Ownership is constrained by observation conditions: students may rehearse defenses, generate evaluative statements with AI, or perform well on a familiar variation without retaining adaptable understanding.
- The four-domain structure is a design proposal, not an empirically established factor structure, and the scoring weights are an illustrative choice rather than a research-supported optimum.
- Local adaptation and validation are required before the framework is used for consequential certification, and disciplines differ in their relevant evidence and acceptable resources.
Citation
Rothwell, W. J. (2026). Assessing Learning When AI Can Produce the Product: An Integrated Framework for Cognitive Ownership and Mutual Evaluation. International Journal of AI in Pedagogy, 2(1). https://doi.org/10.46787/ijaipil.v2i1.8053