Research Article
Assessing teachers' AI literacy: a systematic review of measurement tools
Synthesis: Zainal, Mohd Matore, and Maat catalog the instruments used to measure teacher AI literacy, searching four databases for empirical instrument development published between 2019 and 2025 and appraising 33 studies against a bespoke framework adapted from COSMIN and Terwee et al. (2007). The headline pattern is what the authors call methodological monotony: 31 of 33 instruments (93.9%) are self-report scales of perceived confidence, only two (6.1%) test knowledge objectively, and no performance-based tasks were found. Measurement quality is lopsided. Internal consistency is the strongest domain (28 of 33, 84.8% at Grade A) while fairness is the weakest, with only five instruments (15.2%) reporting measurement invariance or differential item functioning evidence. Content also lags the technology: 29 instruments (87.9%) target general AI concepts and only four (12.1%), all from 2025, address generative AI. The authors read the field as technically sound but inferentially fragile, and argue for Item Response Theory and performance tasks alongside Self-Assessment to separate validated capability from reported confidence.
Key Findings
- Across 33 instruments, 31 (93.9%) captured teachers' subjective perceptions, and only two (6.1%) used objective knowledge tests; no performance-based tasks appeared between 2019 and 2025.
- Fairness was the weakest domain: only five instruments (15.2%) reached Grade A through invariance testing, while 15 (45.5%) were Grade B and 13 (39.4%) Grade C.
- Content validity rested mostly on qualitative review: 21 instruments (63.6%) reached only Grade B and seven (21.2%) Grade C for lacking quantitative expert agreement statistics.
- Structural validity was strong, with 24 instruments (72.7%) at Grade A through CFA, PLS-SEM, or IRT modeling, though none used IRT or Rasch as primary evidence.
- Generative AI remains a niche target: only four instruments (12.1%), all published in 2025, measured GenAI competencies, against 29 (87.9%) built around general AI concepts.
- Output accelerated sharply: 2024 and 2025 accounted for 81.9% of the studies (n=12, 36.4% and n=15, 45.5%), led by China with 11 studies (33.3%) and Turkey with five.
What the review screened
Following PRISMA guidelines, the authors searched Scopus, Web of Science, ERIC, and IEEE Xplore for records published between January 2019 and August 2025, a window opened before the late-2022 diffusion of generative AI. The databases returned 923 records; deduplication in Mendeley removed 560, leaving 363 for title and abstract screening, where two independent reviewers excluded 285. Of the 78 full texts assessed, inter-rater agreement was substantial at 88.5% (69 of 78) with Cohen's kappa of .80 (95% CI [.68, .92]), and nine articles required a consensus meeting. Full-text review excluded 45 articles: 32 for measuring a different construct, four for non-teacher populations, three for reusing an instrument without revalidation, and six for methodological or quality problems. The final sample was 33 empirical studies.
Which instruments were catalogued
Catalogued tools mix named scales with unnamed batteries and range widely in size, from a 12-item scale developed by Uğraş et al. (2024) to a 51-item instrument by Ferikoğlu and Akgün (2022), averaging roughly 30 items. The 5-point Likert format dominated, used by 25 studies (75.8%); four studies (12.1%) used 7-point scales and two used a 6-point forced-choice scale, while 20 studies (60.6%) administered online and 11 (33.3%) did not report the mode. Dimensions cluster into recurring domains. AI knowledge and understanding was most frequent, in 28 of 33 instruments (84.8%). Pedagogy was explicit in 18 studies (54.5%), eight built on TPACK and 10 on alternative constructs. Ethical and social awareness appeared in 20 (60.6%), technical proficiency in 14 (42.4%).
How well the instruments perform
Quality was graded A, B, or C across four domains using a decision matrix adapted from COSMIN and Terwee et al. (2007), benchmarked within each analytical tradition so studies were judged against their own fit standards. Content validity was uneven: five instruments (15.2%) earned Grade A for quantitative consensus evidence such as CVI, Kappa, or IIOC, 21 (63.6%) Grade B, and seven (21.2%) Grade C. Structural validity was stronger, with 24 (72.7%) at Grade A, six (18.2%) at Grade B on exploratory factor analysis alone, and three (9.1%) at Grade C. Internal consistency led at 28 Grade A (84.8%), three Grade B (9.1%), and two Grade C (6.1%). Fairness lagged at five Grade A (15.2%), 15 Grade B (45.5%), and 13 Grade C (39.4%).
What this means for practice
- Faculty developers. Read self-report AI literacy scores as confidence, not capability: 31 of 33 instruments assessed perceived knowledge, so scores cannot stand in for competence.
- Faculty developers. Check fairness evidence before comparing groups, since only five instruments (15.2%) reported formal measurement equivalence.
- Researchers. Ask what a tool targets and how it was validated: 29 of 33 instruments (87.9%) cover general AI concepts, and 21 (63.6%) reported no quantitative expert agreement evidence.
Limitations
- Only English-language records from four databases were searched, so instruments validated in other languages, including work from China and Turkey (16 instruments, 48.5%), were missed.
- Gray literature such as dissertations and conference proceedings was excluded, which may bias the sample toward established instruments.
- Because 31 of 33 instruments were self-report, the synthesis describes how AI literacy is conceptualized and measured, not how it is enacted in practice.
- The appraisal was conducted by the first author using a pre-specified decision matrix, and the review captures a snapshot of a fast-moving field.
Connected Concepts
- AI Literacy
- Teacher AI Competency
- Teaching
- Educational Measurement
- Self-Report Measures
- Self-Assessment
- Assessment Validity
- Psychometrically Aware AI
- Item Response Theory
- Workplace Learning
Connected Articles
- How to Assess AI Literacy: Misalignment Between Self-Reported and Objective-Based Measures — Parallel self-report and objective measures of teacher AI literacy barely agreed across 288 K-12 teachers.
- The GenAI Skill Bypass: Mapping Divergent Pathways of University Students and Staff AI Literacy — Divergent AI literacy pathways among university students and staff.
- Raising Ethical Awareness of GenAI Use Through Student Self-Assessment in the Transition to Higher Education — Guided GenAI self-assessment builds ethical awareness in transitioning higher education students.
- Technology, Education and Critical Media Literacy: Potential, Challenges, and Opportunities — Critical media literacy as a frame for evaluating AI-mediated information.
- Measuring Acceptance of Age-Tiered AI Literacy Guidebooks: A Developmentally Informed Study of K-12 Students and Teachers — Measuring acceptance of developmentally tiered AI literacy guidebooks across K-12 students and teachers.
- Measuring Artificial Intelligence Literacy: A Systematic Review of Instrument Development, Conceptual Foundations, and Psychometric Quality — PRISMA review of 47 AI literacy instruments and their psychometric quality.
Citation
Zainal, M. A., Mohd Matore, M. E. E., & Maat, S. M. (2026). Assessing teachers' AI literacy: a systematic review of measurement tools. Interactive Learning Environments.