On this page

Synthesis: This study develops and validates the Generative AI Assessment Literacy Scale (GAA-LS), an 18-item measure of how well higher education students can use generative AI responsibly in assessed work. Validation across two samples supports its five-dimension structure, its reliability, and its construct validity. GAA-LS scores rise with feedback engagement, academic integrity intention, and responsible AI use intention, and the association with academic integrity intention runs partly through feedback engagement.

Key Findings

  • The scale defines GenAI assessment literacy as five related judgments. Assessment criteria awareness, AI–task appropriateness judgment, verification and evidence checking, ethical attribution and academic integrity, and feedback uptake and revision literacy, each measured by a short Likert-scale subscale.
  • It is an integrative construct rather than a replacement for existing literacies. The scale targets the point where general AI Literacy, Assessment literacy, Feedback literacy, and Academic Integrity are jointly required: judging whether AI support fits the task, checking whether AI output can be trusted, disclosing assistance, and revising work without surrendering authorship.
  • Reliability was strong. Internal consistency was satisfactory to strong across every subscale and for the total scale, and composite reliability and average variance extracted supported the measurement model.
  • The five-factor structure held, with broad validity support. The hypothesised model fitted an independent confirmatory sample well and clearly outperformed higher-order, three-factor, and single-factor alternatives, with an ordinal-estimator check reaching the same conclusion. Convergent, discriminant, criterion-related, and known-group validity evidence were all obtained.
  • The scale behaves comparably across student groups. Measurement invariance held across gender, discipline, and AI-use frequency, so the instrument can be used to compare these groups with reasonable confidence.
  • Higher literacy traveled with more feedback engagement and stronger integrity intentions. Total scores correlated positively with feedback engagement, academic integrity intention, and responsible AI use intention, most strongly with feedback engagement. Subscale patterns were theoretically coherent — the feedback-uptake dimension tracked feedback engagement most closely, and the ethical-attribution dimension tracked academic integrity intention most closely.
  • Feedback engagement partly explains the integrity link. In the structural model, GenAI assessment literacy was associated with academic integrity intention both directly and indirectly through feedback engagement. Because the data are cross-sectional, this is evidence of indirect association rather than causal mediation.
  • Experience and instructional support accompany stronger literacy. Students reporting frequent AI use, prior AI training, or course-level disclosure rules scored higher than their counterparts, by small to moderate margins.

Study Design & Method

The study is a two-study cross-sectional scale development and validation project, following established instrument-development guidance. An initial item pool drawn from the literature on assessment literacy, AI-supported assessment, feedback engagement, and academic integrity was refined through review by an eight-member expert panel, cognitive interviews and pilot testing with students, and exploratory factor analysis, leaving a final version of 18 items with five dimensions and at least three items per dimension. Every item is written to refer explicitly to course assignments, rubrics, feedback, disclosure rules, or revision decisions; no GenAI platform was evaluated by name, and no AI system generated responses or made psychometric decisions.

Questionnaires were returned by students at six higher education institutions in China, and 1,284 valid responses were retained and randomly split. The exploratory subsample (486) carried the item analysis and exploratory factor analysis, run as principal axis factoring with oblimin rotation. The confirmatory subsample (798) carried confirmatory factor analysis and model comparison, reliability testing, convergent and discriminant validity, measurement invariance testing across gender, discipline, and AI-use frequency, criterion-related and known-group validity checks, and structural equation modeling in which the indirect association through feedback engagement was estimated by bootstrapping. Alongside the GAA-LS, participants completed short measures of feedback engagement, academic integrity intention, and responsible AI use intention. Common-method bias was examined through procedural design plus Harman's single-factor test, a one-factor confirmatory comparison, and a common-latent-factor sensitivity check, treated as diagnostic rather than conclusive.

Implications

  • The scale is a diagnostic, not a verdict. Its five subscales let instructors locate where a class is weak — disclosure rules, evidence checking, criteria interpretation, or feedback uptake — and teach to that gap instead of issuing blanket warnings about AI.
  • Assessment briefs should make the expectation set explicit. Criteria, permitted and prohibited AI support, expectations for checking AI output, acknowledgment rules, and how feedback should feed revision all need to be visible to students, and the scale names exactly those elements.
  • Institutions gain a way to test whether policy becomes usable knowledge. The instrument can be used when introducing disclosure rules or redesigning tasks to check whether students actually understand how to comply, rather than assuming that publishing a rule is enough.
  • The framing connects assessment literacy, feedback literacy, and academic integrity in one framework. Its grounding in Self-Regulated Learning and evaluative judgment links understanding criteria, monitoring the quality of one's own work, and acting responsibly on feedback, which is where AI-supported assessment in Higher Education is most exposed.

Limitations

  • Cross-sectional and self-reported, so the associations are not causal; longitudinal and experimental work is needed to test whether GenAI assessment literacy predicts later feedback behavior, disclosure, and assessment outcomes.
  • Self-report measures are open to social desirability and common-method variance; the statistical checks found no dominant method factor, but they cannot rule bias out.
  • Academic integrity intention may not fully predict actual behavior, and the ethical-attribution subscale is conceptually close to that criterion, so those correlations are better read as proximal validity evidence than as proof of construct separation.
  • The sample came from one national context and was not designed to be nationally representative; cross-cultural and cross-institution invariance, test-retest reliability, and invariance across separately sampled undergraduate and postgraduate groups remain untested.
  • The validation did not include established AI literacy or assessment literacy instruments, so incremental validity beyond those adjacent constructs is still unknown; behavioral indicators such as revision logs or disclosure statements would strengthen ecological validity.

Connected Concepts

Connected Articles

Citation

Nie, J., Zhang, Z., Lu, X., Zhang, Y., & Zhang, M. (2026). Development and Validation of the Generative AI Assessment Literacy Scale for Higher Education Students: Psychometric Evidence and Associations with Feedback Engagement and Academic Integrity. Frontiers in Education, 11, 1934632.

Embed this page

Copy the code below to embed a chromeless version of this page in a learning management system or other website. The embedded view hides the site header, navigation, and footer.