Research Article
How AI-vulnerable is Australian higher education assessment? A computational audit of Group of Eight Arts and Humanities units, 2022–2026
Synthesis: Universities have spent four years writing GenAI policy while their published assessments have barely moved. Villanueva audits 15,587 annual unit records and 53,915 assessment items across Australia's Group of Eight Arts and Humanities units from 2022 to 2026, scoring each item's exposure with an author-developed AI Vulnerability Index. Average unit exposure rose sharply at all seven institutions with comparable data — 0.269–0.330 in 2022 to 0.559–0.673 in 2024 — and every one stayed above its 2022 level in 2026. Exposure is concentrated rather than diffuse: five unsupervised written formats carry nearly all of it, and take-home essays plus research reports alone account for 71.4% of exposure among items scoring at least 0.60 once weighted by marks. High exposure is neither misconduct nor automatically a defect, but where it is unintended it weakens the evidential basis of the credentials a program awards.
Key Findings
- The audit covers 15,587 annual unit records and 53,915 assessment items at the eight Group of Eight universities from 2022 to 2026, collected from public handbooks, course profiles and subject outlines. Units are counted separately in every year they were recorded, so the total is not a count of distinct units.
- Across the five years the average item score is 0.505; 47.3% of items score at least 0.60 and 34.8% at least 0.80. In 2026, 58.6% of items were highly or very highly exposed and 56.5% very highly exposed, the highest yearly figures in the series.
- Average unit exposure rose at every institution with comparable 2022 and 2024 data, from 0.269–0.330 to 0.559–0.673, and all seven remained above their 2022 level in 2026, though later trends diverged: UNSW kept rising, several institutions changed little after 2024, and Queensland declined from its 2024 peak.
- In 2026, very-high-exposure items absorbed between 53.1% (Queensland) and 70.7% (UNSW) of allocated marks, and between 53.1% and 79.1% of units scored at least 0.60.
- Exposure concentrates in five formats: take-home essays (34.7% of items scoring at least 0.60), research reports and literature reviews (32.8%), unproctored online quizzes (13.1%), reflective journals and essays (7.4%), and unsupervised creative writing or digital artifacts (6.8%). Weighted by contribution to final grades, the first two alone account for 71.4%.
- Institutional policy category did not correspond to lower exposure: mandate and framework institutions both appear at the higher and middle of the distribution, so the prescriptiveness of central policy does not reveal the conditions under which students produce their work.
The AI Vulnerability Index: exposure, not misconduct
The AIVI scores each assessment item from 0.00 to 1.00 using published task format, supervision, identity checks, evidence of how work developed, requirements tied to local material, and the author's evaluation of GenAI capability in that year. Scores are banded as low (below 0.30), moderate (0.30–0.59), high (0.60–0.79) and very high (0.80 or above), and each item's score is multiplied by its share of the final grade before unit and university averages are taken — an 80% essay with a 20% exam scores far higher than the reverse. Scores are year-relative, so an unchanged task can score higher later; the index measures how readily GenAI available in a given year could help complete the work, not whether students used it. A high score may describe a task deliberately designed to permit AI, and exposure is separate from security: the secure/non-secure arrangements in sector guidance govern who completed the work and under what conditions, while AIVI governs how easily it could be assisted. Validation is substantial but self-referential. Five open-weight models classified the items; agreement across all ten model pairs exceeded 0.80 for 2022–2023, with an overall weighted kappa of 0.907 across 233,624 comparisons and an average of 0.894 for 2024–2026. Human review of 326 checks judged 315 correct or mostly correct (96.6%). Both checks assess extraction and classification, not independent calibration of the numerical ratings.
What changed, and which formats carry it
The low-exposure end of the scale is instructive. Invigilated examinations score 0.00, oral examinations and vivas 0.05, live performance and laboratory work 0.10, supervised in-class tests 0.15, and proctored online examinations about 0.40. At the other end, take-home essays, unproctored online quizzes, reflective journals, research reports and unsupervised digital artifacts score at or near 1.00 — and recorded presentations and remote group projects reach 0.85 despite their oral or interactive form, because the public record cannot verify how the work was produced. Exposure therefore depends on the conditions in which work is completed rather than on educational purpose alone, and tasks with the same learning aims can have very different scores. This is where the audit's findings bite hardest for authentic assessment. Formats with an authentic surface — recorded presentations, digital artifacts, reflective journals — sit in the high-exposure band under 2026 scoring, bearing out the warning that authentic design does not by itself prevent misconduct or verify authorship. The concentration matters for practice: a small number of established, unsupervised written formats accounts for most institutional exposure, so the priority list for redesign is shorter than much of the post-ChatGPT literature implies. That list is also harder to act on in the Arts and Humanities than elsewhere, because there the essay is not an instrument documenting learning achieved elsewhere but the medium through which critical reasoning, sustained reading and independent research are developed. Replacing a research essay with a viva or a supervised in-class task is not an equivalent swap.
Policy did not predict exposure
Institutional documents were classified as mandate (specified steps to verify learning or redesign assessment), framework (structured principles leaving implementation to faculties) or guidance (advisory). Mandate and framework institutions appear across the distribution of unit exposure and of mark allocation, so more prescriptive central policy did not consistently accompany less exposed assessment. The audit's reading is that policy and redesign are different processes: a university may publish a sophisticated position on GenAI use, disclosure and detection while unit coordinators work within inherited templates, approval processes, workload constraints and uneven local support. Reform is an organizational problem before a pedagogical one — moving from unsupervised essays to vivas, staged checkpoints or supervised production requires rooms, invigilation, moderation, accessibility adjustments, staff time, and smaller marking loads, and large-scale online proctoring offers an apparent shortcut at documented ethical and equity cost. The author's constructive proposal is sufficient assurance at program level: map each task to its learning outcomes, intended GenAI role, contribution to the final grade and arrangements for verifying the work, and make sure the mix follows from the outcomes being certified rather than from inherited templates. Audits should be repeated, since a defensible mix in 2026 may not yield the same evidence in 2030.
What this means for practice
- Assessment designers. Audit the mark-weighted mix rather than counting items: a low-weight oral check can verify learning demonstrated in a much larger open task, and exposure measured by marks is what determines how much of a grade rests on unverified work.
- Institutions and administrators. Treat policy text as a statement of intent, not evidence of redesign. The binding constraints are organizational — rooms, invigilation, moderation, accessibility adjustments and staff time — so program-level coordination has to be funded rather than assumed to follow from guidance.
- Instructors. Read a high exposure score as a prompt for a design decision, not an automatic instruction to move to supervised assessment. The paper's own position is that open assessment remains valuable and that heavy reliance on invigilated examinations narrows what students can demonstrate.
- Researchers. Note how little headroom the index now has: most unsupervised written formats already score at or near the 2026 maximum, so further capability gains will barely move the numbers. Future audits need to ask what evidence unsupervised written work can provide when advanced systems can approximate the submitted product.
Limitations
- The AIVI is a research-informed exposure rating, not a measure of AI use, misconduct or educational value; its ratings and category thresholds rest on evaluative judgment, and different choices would change the reported averages and proportions.
- Only public assessment information was used. Instructions on learning management systems, unpublished rubrics, oral checks, draft requirements and local GenAI rules are invisible, and where a published description does not match local delivery — a quiz recorded as online but administered in tutorials — the classifier defaults to the public text.
- Classification by format misses design nuance: a research essay with staged drafts, local data or an oral defense may be considerably less exposed than a generic take-home essay, and the scheme cannot see the difference.
- Year-relative scoring means the index cannot separate rising capability ratings from assessment redesign, and because unsupervised written formats already sit near the maximum in 2026, it has little room to register further change in the formats that matter most.
Citation
Villanueva, G. (2026). How AI-vulnerable is Australian higher education assessment? A computational audit of Group of Eight Arts and Humanities units, 2022–2026. Assessment & Evaluation in Higher Education.