Research Article
Mapping the Authorized Boundary: A Comparative Policy-Vignette Study of Generative AI Governance in Australian Higher Education
Synthesis: Australian universities regulate student use of Generative AI through a layered mix of binding instruments and official guidance, yet identical conduct can be classified differently across institutions. This study applied 15 standardized vignettes to the public policy environments of 20 Australian universities, producing 300 university-case classifications. Each combination was coded from clearly permitted to indeterminate, and cross-university divergence was measured with normalized Shannon entropy. No combination met the strict threshold for clearly permitted; 120 (40.0%) were clearly prohibited and 56 (18.7%) remained indeterminate. Binding instruments stayed silent on GenAI in 100 combinations, and guidance resolved 88 of them. An independent human second coder agreed with 57.3% of the primary classifications (Cohen's kappa = 0.395). Written AI Governance supports a defensible classification only when it states a default, names who may vary it, and specifies a fallback — a question of Educational AI Policy design rather than of permissiveness alone.
Key Findings
- Across 300 university-case combinations, 120 (40.0%) were classified as clearly prohibited, 97 (32.3%) as potential policy breaches, 27 (9.0%) as permitted with conditions, and 56 (18.7%) as indeterminate; none reached clearly permitted.
- Disclosed language rewriting and a disclosed AI-drafted paragraph produced the highest divergence (H = 0.801), while an explicit assessment prohibition produced complete agreement across all 20 universities (H = 0.000).
- Binding instruments remained silent on GenAI in 100 combinations and guidance clarified 88 of them, so Layer A alone or agreement between layers resolved only 188 of 300 classifications (62.7%).
- Process-only assistance produced more indeterminacy than retained AI text: each outline case yielded 11 indeterminate classifications and verified source discovery 16, while no major-drafting case exceeded one.
- Disclosure did not determine substantive outcomes: 117 requirements were satisfied, 80 were false declarations, and the same disclosure status still produced different classifications depending on permission rules and task instructions.
- An independent human second coder working without AI assistance matched 43 of 75 primary classifications (57.3%; unweighted Cohen's kappa = 0.395), indicating only moderate agreement on the five-category nominal scale.
Layered authority and the missing default
Coding separated binding policies, procedures, and misconduct provisions (Layer A) from the full official environment (Layer B), which adds guidance for students, staff, and assessment. Layer B supplied the operative rule in 109 of 300 combinations (36.3%), and guidance clarified 88 of the 100 combinations where binding instruments never mentioned GenAI. Silence in the binding layer did not predict ambiguity. Curtin, Tasmania, and RMIT all had silent binding layers across all 15 cases, yet their indeterminacy rates were 6.7%, 26.7%, and 66.7% respectively, because Curtin's guidance supplied a written-permission default while RMIT pointed students to course requirements with no fallback. Educator discretion appeared in 204 combinations (68.0%) and an express prior-permission requirement in 128 (42.7%), while policy conflict was recorded in 24 (8.0%). Discretion is not itself the defect; unwritten discretion with no stated holder, scope, or fallback is.
Two competing governance architectures
Institutions divided into permission-gated systems, which treat silence as denial until a named authority grants approval, and disclosure-based systems, which permit some uses by default and then impose limits based on contribution or harm. That contrast explains the two most divergent vignettes, both moderate-contribution cases with H = 0.801. It also explains why removing disclosure changed so little: 14 of 20 classifications were unchanged for rewriting and paragraph generation, and 16 for major drafting. Disclosure cannot cure unauthorized use in a permission-gated system, nor prohibited substitution of generated work in a disclosure-based system, and a false declaration can only aggravate conduct that already occupies a restrictive category. Universities must therefore say whether AI Use and Disclosure Statements records use, gates permission, or evidences authorship.
Retained text, verification, and task authority
The written environments favored visible outputs. Rules on authorship, editing, and unauthorized assistance captured language rewriting and retained generated paragraphs, but planning and verified source discovery were less often addressed — a theme carried by 80 supporting rows. Verification failure and explicit task rules carried 40 rows and produced the strongest shifts: between verified source discovery and fabricated references, 15 universities moved from indeterminate to a determinate outcome, and between a silent instruction and an explicit prohibition, 11 cases were resolved from indeterminacy. Case 15 was clearest: all 20 universities treated an explicit assessment ban as controlling. Established misconduct categories — contract cheating, ghostwriting, fabrication, misrepresentation — resolved 29 supporting rows of evidence without any new technology-specific rule.
What this means for practice
- Instructors. Write the task rule into every assessment instruction: state whether GenAI is permitted, limited, or prohibited, and what applies when the instruction is silent — explicit prohibitions produced unanimity, moderate contribution cases diverged most.
- Policy authors. Separate disclosure from authorization, and state which function acknowledgment serves: recording use, gating permission, evidencing authorship, or several of these at once.
- Curriculum and assessment designers. Cover the full continuum from planning and source discovery to rewriting and retained generated text, and pair permission language with an express duty to verify claims, quotations, and references.
- Quality assurance staff and regulators. Run policy-vignette stress tests that ask a provider to cite the document, authority, and fallback behind each classification, and repeat the exercise after revisions.
- Institutional governance leads. Publish both a fallback default and a conflict hierarchy that says which instrument controls when policy, guidance, and task instructions disagree.
Limitations
- The purposive sample of 20 institutions does not statistically represent the Australian sector, and the 150 indexed documents (138 included: 59 binding instruments and 79 guidance documents) were collected on July 21-22, 2026, so they may since have changed.
- Fixed vignettes simplify conduct and omit sanctions, evidentiary disputes, procedure, and local context; the study analyzes public written documents, not decisions or enforcement.
- The independent human second coder achieved only moderate agreement with the human-led, AI-assisted primary coding (57.3%; kappa = 0.395), with 32 disagreements clustering at the potential-breach/clearly-prohibited and indeterminate/permitted-with-conditions boundaries.
- Normalized Shannon entropy measures category dispersion and cannot distinguish desirable governance diversity from problematic inconsistency.
Citation
Poudyal, B. (2026). Mapping the Authorized Boundary: A Comparative Policy-Vignette Study of Generative AI Governance in Australian Higher Education. arXiv preprint.