Research Article
What Does the Credential Still Certify? Cognitive Stewardship for AI-Mediated Education
Synthesis: Generative AI undermines a basic premise of educational assessment: that submitted work reliably evidences the human capacities a credential certifies. Yao (2026) develops cognitive stewardship, a framework linking four elements — the learning claim, the delegation boundary, the evidence standard, and safeguards — to reason about what remains inferable about learning once cognitive work is delegated to AI. The paper then audits verified public GenAI assessment guidance from 30 universities across five English-speaking systems, finding that institutions are getting better at classifying AI use than at explaining what evidence of learning remains valid under each class.
Key Findings
- The boundary–evidence asymmetry. Across 30 audited university policy packages, the mean delegation-boundary score was 2.47/4 but the mean evidence-standard score was only 1.89/4 — 22 policies scored higher on boundary than evidence, 3 tied, and only 5 reversed the gap. Public guidance draws lines around what AI may do more often than it explains what evidence of learning must remain.
- Safeguards are present but sparse. Packages contained a mean of 2.75 of 8 possible safeguards: privacy was most visible (67%), followed by detection caution (47%), Accessibility (42%), appeal/due process (33%), tool-access equity (32%), vendor governance (29%), non-AI alternatives (22%), and workload/proportionality (4%). Institutions ask for disclosure more often than they provide protection, recourse, or alternatives.
- Policies are clearest for final-output substitution. Substitution averaged 2.49/4 actionability with clear answers in 83% of packages, while the other five scenarios (access, Feedback support, process, output-verification, programming workflow) ranged from 1.17 to 2.32. Guidance is strongest when AI use resembles cheating and weakest when it resembles learning support or authentic professional workflow.
- Output-verification support is the thinnest scenario. Only 23% of packages directly covered the capacity to inspect, challenge, and correct AI output — arguably one of the strongest reasons to teach with AI. Only 31% required evidence connecting access support back to the assessed claim.
- Four policy archetypes. Nine packages were higher-clarity (connecting permission to evidence and safeguards), four were boundary-forward (visible AI-use categories with thin evidence/safeguards), eight had limited evidence visibility, and nine were partially connected. The framework identifies a design gradient, not a compliant/noncompliant binary.
The educational delegation problem
The paper names the core problem educational delegation: not whether AI touched the work, but which cognitive operations moved from the learner to the system and which remained. One student may use AI feedback while retaining problem formulation, source evaluation, revision judgment, and final responsibility; another may delegate topic selection, evidence search, argument structure, drafting, and citation. Both involve AI, but they support very different educational inferences.
Assessment validity asks whether a task produces evidence for a learning claim; credential validity asks whether accumulated evidence justifies an institutional claim about the learner. Artifact-centered Assessment is fragile because a final product can look excellent while revealing little about which operations the learner performed — and detection does not solve this, since even perfect AI Detection would not show whether the use was educationally appropriate or displaced the target capacity.
The cognitive stewardship framework
Cognitive stewardship links four elements:
- The learning claim being certified
- The delegation boundary specifying what AI may do
- The evidence standard showing what observable work, explanation, verification, or defense remains
- The safeguards protecting privacy, accessibility, proportionality, equity, and appeal
Its unit of analysis is not an individual competency checklist but the warrant behind a course, program, or credential. A policy is under-specified when any element is missing — a course can clearly permit AI and still fail to say what it can certify, or prohibit AI and still fail to protect access or due process. The main design rule: start with the certified claim, not the tool. If delegation would remove the operation being certified, the boundary should be restrictive or the task redesigned; if delegation supports the claim but changes the evidence, the policy should add process evidence, explanation, verification, or defense.
The framework uses five recurring delegation types: access support (translation, speech-to-text, assistive Scaffolding), feedback support (critique while the learner retains judgment), process support (planning, drafting, debugging), substitution (AI performs the operation being assessed), and output-verification support (treating checking AI-mediated work as a taught, assessed capability). These are role descriptions, not moral labels — the same use can be access support in one task and substitution in another.
Audit method
The audit scored public institutional policy packages (the set of official sources through which each university tells students how GenAI may be used in assessed work) across the UK (11), Australia (5), New Zealand (2), Canada (5), and the US (7). A pre-specified, source-grounded scoring codebook was applied by four open-weight LLMs as structured coders, with scores averaged to dampen single-model bias. The audit treated public guidance as a reader-facing artifact — what a student, instructor, or reviewer can see about claim, boundary, evidence, and safeguard. Between-model agreement was a sensitivity measure, not validation, so exact score levels are exploratory descriptions.
What this means for practice
- Administrators. Start from the certified claim rather than the tool: name the cognitive operations a credential must still evidence, then write the delegation boundary from that claim, which shifts Assessment Validity work away from detection and Academic Integrity enforcement.
- Administrators. Require every AI-use category to state what evidence stays valid under it: across the 30 audited packages, delegation boundaries averaged 2.47/4 while evidence standards averaged only 1.89/4, and only 23% of packages covered inspecting, challenging and correcting AI output.
- Faculty developers. Publish scenarios and worked examples for the uses your guidance handles worst, since substitution was clear in 83% of packages while process support, Feedback, and programming workflow ranged from 1.17 to 2.32 out of 4.
- Administrators. Treat Privacy, Accessibility, due process, non-AI alternatives and workload as conditions of the warrant rather than exceptions: packages averaged only 2.75 of 8 safeguards and workload/proportionality appeared in 4%; disclosure and monitoring can make students more visible to the institution without making assessment fairer, and surveillance harms fall unevenly on racialized, disabled, low-income, international and linguistically marginalized learners.
- Researchers. Do not anchor policy in today's model weaknesses such as hallucination or bias; ask which human capacities must remain visible even when task performance can be delegated, and treat stewardship as a AI Governance arrangement rather than an assessment technique.
Limitations
- The corpus is purposive, not representative: 30 public policy packages from five English-speaking systems (United Kingdom 11, Australia 5, New Zealand 2, Canada 5, United States 7), assembled to compare visible policy designs rather than to estimate worldwide prevalence.
- Scores came from four open-weight LLMs applying a single codebook with no independently human-coded comparison set, so between-model agreement is a sensitivity measure rather than validation: mean standard deviation was 0.57 for delegation-boundary scores, 0.53 for evidence-standard scores and 0.42 for learning-claim scores, with high-variation cases appearing in 13 to 22 of the 30 packages depending on the scenario.
- The audit measures published policy text, not classroom practice, institutional intention or learning outcomes, and the constructs sit on different scale maxima (learning claim 0-3; boundary and evidence 0-4), so exact score levels are exploratory descriptions.
Citation
Yao, K. (2026). What Does the Credential Still Certify? Cognitive Stewardship for AI-Mediated Education. (cs.CY). Accepted at AIES 2026.