Research Article
A bit of chaos and madness: The AI Assessment Scale and the work of assessment reform
Synthesis: This study examines the implementation of the Artificial Intelligence Assessment Scale (AIAS), a structured framework for redesigning university assessment in response to generative AI. Surveying 30 academic staff, the researchers found that while the framework's transparency and guidance were valued, implementation was hampered by departmental inconsistencies, workload pressures, and uncertainty about appropriate AI use levels.
The study frames assessment reform through the lens of 'assessment security', conceptualizing the work of redesign as a security practice in higher education. Staff described the process as 'a bit of chaos and madness', capturing both the disruptive potential and the lack of coordinated institutional support. This connects to broader challenges around academic integrity in the age of generative AI.
The findings have implications for faculty development and institutional policy, suggesting that successful AI assessment frameworks require not just clear guidelines but also adequate resourcing, departmental alignment, and ongoing professional support. The study contributes to the growing literature on implementation science for AI literacy in higher education.
What this means for practice
- Instructors. Write the task and the AIAS level together instead of stamping a level onto an existing assignment: staff read the levels as permissions for AI use rather than specifications of the assessable work, and the scale degraded into a compliance exercise when retrofitted.
- Instructors. Ask what students must produce at each level — which reasoning, evidence, or justification has to remain theirs — so the level constrains the task rather than licensing a shortcut.
- Instructors. Budget realistically for redesign, because implementation was hampered by workload pressures and uncertainty about what levels of AI use were acceptable, not by the framework's wording.
- Instructors. Secure departmental alignment before adopting a level, since inconsistent practice across departments and a lack of coordinated institutional support produced the "bit of chaos and madness" staff described.
- Instructors. Pair the scale with explicit critical AI literacy outcomes, so students learn to evaluate AI output rather than only to be told how much AI they may use.
Limitations
- The evidence comes from 30 academic staff in five focus groups at two institutions (n = 17 in three groups at a private international university in Vietnam; n = 13 in two groups at a single UK business school), and the authors explicitly make no causal or comparative claims between sites.
- Sampling differed by institution — purposive at one, self-selection at the other — which limits comparability and under-represents disciplines outside the participating school.
- Primary coding was done by one member of the research team, so no inter-rater statistics were available; the authors mitigated this with an external reviewer and cross-team discussion of divergent codes.
- Focus groups are shaped by group dynamics and social desirability, which the authors note constrains what staff disclose about their assessment practice.
Citation
Mike Perkins, Darius Postma, Jasper Roe, Susan Sisay, Craig Holdcroft (2026). 'A bit of chaos and madness': The AI Assessment Scale and the work of assessment reform. arXiv cs.HC.