On this page

AI literacy — the knowledge, skills, and critical dispositions needed to understand, evaluate, and effectively use AI Technologies in educational contexts. AI literacy spans foundational understanding of how AI works, practical competence in using AI tools, critical evaluation of AI outputs, and ethical awareness of AI's societal implications.

Questions to Consider

  • What does it mean to be 'AI literate'? Is it knowing how to use the tools, understanding how they work, or being able to critically evaluate their output — and which matters most for your own goals?
  • Self-reported AI literacy diverges sharply from measured performance — teachers overestimate their skills by about 40%. How confident are you in your own AI skills, and what evidence would you trust to actually test that confidence?
  • AI literacy is described as a metacognitive social practice, not a checklist of skills. Because LLMs are probabilistic and opaque, what does it mean to 'critically evaluate' output when you can't see inside the model?
  • There's a tension between critical-use literacy for academic settings and workflow-integration literacy for employment — educators and employers value different skills. Which kind of AI literacy is your context actually teaching?
  • Technical AI-literacy training alone, without self-efficacy and self-regulation support, can actually increase dependency on AI. How could learning to use AI make you more reliant on it rather than more capable?
  • AI literacy is also a recognition skill: spotting when an AI is agreeing with you because it's being sycophantic versus because you're right. Have you ever caught an AI just telling you what you wanted to hear?

Introduction

AI literacy has rapidly emerged as a core competency for learners, educators, and institutions as Generative AI becomes embedded in education. Unlike general digital literacy, AI literacy requires understanding probabilistic systems that can hallucinate, exhibit bias, and shift Learner Agency from human to machine — making Critical Thinking and Over-Reliance central to the construct.

Dimensions of AI literacy

AI literacy research in this knowledge base spans four interconnected dimensions:

Frameworks increasingly trace how these dimensions are enacted in practice: Dai & Chan (2026) show how postgraduate researchers enact all four dimensions when using Generative AI across the research workflow and scaffold research-specific responsible-use guidelines from them, while Şan & Orhan Karsak (2026) use Word-Association mapping to show that Turkish undergraduates' AI cognition is instrumentally "black-box" dominated, with ethical-governance concepts structurally isolated (τ=−0.819) — evidence that credential design must deliberately bridge experiential tool use and ethical frameworks.

Foundational knowledge: Understanding what LLMs are, how they differ from rule-based systems, and their fundamental limitations. This includes awareness of model capabilities, training data biases, and the distinction between task-specific AI and general-purpose models. Research in Prompt Engineering examines how understanding prompt mechanisms affects effective AI use.

Practical competence: The ability to use AI tools effectively — from Prompt Engineering to interpreting outputs. Studies of student GenAI usage patterns reveal that tool access alone doesn't produce competence; structured practice and Scaffolding are essential. The vibe coding framework shows how K-12 teachers can develop practical AI literacy through guided tool creation. Miles, Haber-Curran and Arar (2026) argue that this dimension has a part optimization training misses: prompt literacy, as distinct from prompt engineering, treats prompting as a rhetorical and ethical act, so instruction that only tunes outputs for better performance leaves the ethical and epistemological dimensions of LLM use untouched.

Critical evaluation: The capacity to assess AI outputs for accuracy, bias, and appropriateness. Research on literacy assessment shows a 40% gap between self-reported and performance-based AI literacy — people consistently overestimate their evaluation skills. That divergence is the clearest case in the knowledge base of a self-report measure failing to track the competence it names. This connects to Over-Reliance research showing that students who Trust AI uncritically learn less. Recent conceptual work pushes evaluation toward verification: the PEARLS framework (Wang) treats AI output as a provisional knowledge claim whose warrant must be assembled and examined across six dimensions (Process, Evidence, Access, Reproducibility, Legitimacy, Source), and advances verification-driven learning as the mechanism by which learners build expertise while checking AI claims. Complementing this, the student-centered responsible-use framework (Alsammani) offers ten student-facing guidelines across Learning and Growth, Ethics and Integrity, and Awareness and Safety pillars that externalize metacognitive prompts at the point of decision. For secondary learners, the AI-Assisted Research Competency (AARC) framework (Beau, Flaquière & Lazar 2026) grounds this in virtue epistemology and AI intuition, defining research literacy as conducting inquiry with AI without surrendering authorship, judgment, verification, or responsibility — operationalized as an analytic rubric and a verify–cite–reflect commitment routine. At the Assessment end, Saleh (2026) proposes a human capability test that turns evaluation into a design principle: ask what a student must demonstrate independently, what can be strengthened through AI augmentation, and what the student must verify, defend, and take responsibility for.

Disciplinary verification raises the same bar further. Vega-Baudrit and Rivera Alvarez (2026) argue that in university chemistry the core AI literacy demand is representational translation across Johnstone's macroscopic, submicroscopic and symbolic domains, because an output can be locally persuasive in one domain while contradicting charge balance, a mechanism or a safety limit in another. Their remedy is verification-centered integration: make verification an assessed activity, keep prompt logs and revision histories as reasoning traces, and treat prompting as an epistemic act that specifies constraints, assumptions and adequacy criteria, with students identifying a false assumption or justifying rejection of a generated answer.

Ethical and institutional awareness: Understanding AI's broader implications — from Academic Integrity to Equity to Privacy. AI literacy at the institutional level involves policy development, Educational Development, and governance frameworks — institutional AI literacy is a matter of policy as much as pedagogy. The EPIQ-AI framework frames institutional AI literacy as a sociotechnical alignment challenge, not just individual training.

  • AI literacy as a governance capacity for sustainable development. Islam, Morshed, and Islam (2026) reconceptualize AI literacy as a governance-oriented capacity rather than a purely educational or technical skill, linking it to all seventeen UN Sustainable Development Goals. Their six-level AIRE Taxonomy (Recognize → Comprehend → Apply → Analyze → Integrate → Govern) extends Bloom's hierarchy by adding ethical synthesis and strategic foresight, positioning advanced competencies (Analyze–Govern) as the pathway from foundational literacy to institutional and policy-level governance — an "18th SDG" heuristic that treats literacy as a cross-cutting cognitive and ethical bridge. A survey of 300 professionals in a national context found strong technical awareness but limited ethical and governance readiness, with ethical reasoning and reflective thinking the strongest predictors of sustainable, trustworthy AI use and governance literacy the strongest predictor of AI–SDG nexus awareness (β = 0.64). This empirically grounds the knowledge base's emphasis on critical-use literacy and ties AI literacy directly to Sustainability and policy integration.

How AI literacy is developed

Research points to collaborative and active approaches as most effective. Rismanchian & Doroudi position AI literacy as more than applied skill: in their AI×Ed framework the learner is a distinct end user of AI (accessed through AI literacy and AI education), and they argue for renewed AI-literacy work that encourages reflection on learning itself — treating literacy as a route back into "AI as an analogy to human intelligence" research on how people learn, a strand the field largely abandoned. The ICAP Framework (Interactive → Constructive → Active → Passive) provides a useful taxonomy of cognitive engagement: students learn AI literacy best when they co-construct knowledge rather than passively receive information. Designers should select the mode that fits the learning goal and favor the deeper (constructive and interactive) modes where possible. Practical activities — designing prompts, evaluating outputs in groups, debating AI Ethics — outperform lectures.

Composition starts from lived experience — but concept depth needs scaffolding. Burriss et al. (2026) found that youth anchored abstract AI Ethics principles in systems they encounter daily — electronic "hall passes," school-regulated laptops, plagiarism detectors — with one group explaining that AI in schools "was personal to all group members," and 15 of 18 end-of-unit survey responses reported that making the video changed their understanding of AI ethics. The authors are explicit about the limit: composing did not by itself guarantee conceptual grounding — "informed consent" was translated as "We gotta approve," and some terms (e.g., "failsafe") were never explained, prompting them to plan deeper conceptual work in later iterations. Creative production develops AI literacy best when explicit conceptual Scaffolding accompanies the making rather than being assumed to emerge from it.

AI intuition as the experiential complement to literacy. A recurring gap is that K-12 frameworks assume learners approach AI through declarative, rule-based knowledge, when in practice they first develop a practical "feel" for how AI responds — experimenting with prompts, observing behavior, adapting strategies — before they can articulate formal principles. Beau & Lazar (2026) formalize this as AI intuition: an experiential, inductive, often tacit understanding that develops through iterative interaction with AI systems and supports context-sensitive judgment under uncertainty. Their dual framework situates AI literacy (structured, largely static competencies) alongside AI intuition (a dynamic learning process), mapping both across the standard dimensions — understand concepts, use tools, evaluate critically, apply ethically, reflect — so learners combine conceptual clarity and Guardrails (literacy) with the practical wisdom to trust-but-verify and decide when to disengage (intuition). Because intuition cultivates "expert observation" of AI (e.g. stress-testing prompts to find where a model fails), the authors argue it is a safeguard against over-reliance and critical-thinking erosion; it is distinct from Prompt Engineering (which optimizes outputs) in foregrounding epistemic judgment. This connects literacy to Experiential Learning, Constructivism, and developmentally grounded classroom practice, and points teacher preparation toward facilitating inductive exploration rather than only delivering concepts.

Motivation is a precondition, not just an outcome. Liang et al. (2026) found, across 2,086 secondary students in a year-long AI curriculum, that students who transitioned into or remained in a Self-Determined motivational profile (high autonomy, competence, and relatedness need satisfaction) showed the greatest AI-literacy gains. AI literacy develops through sustained engagement, and that engagement is itself shaped by motivation and psychological-need support — so effective AI-literacy instruction should attend to learners' motivation, not only their skills.

Instructional emphasis and task openness reach performance through different needs. Yu, Lin and Chen (2026) randomized 320 undergraduates to four groups in a 2 x 2 experiment on an AIGC image-generation task scored with an objective rubric (ICC = 0.89) and found that thinking-based instruction, covering model limitations, critical evaluation and ethics, outperformed skills-based prompting instruction (M = 9.21 vs 7.69, F(1, 316) = 50.79, p < 0.001, partial eta squared = 0.138), with the wider margin on open-ended tasks (10.10 vs 8.00). The structural model shows the two design choices travel by different routes: instruction worked indirectly through autonomy (beta = 0.035) and competence (beta = 0.079), while task openness operated mainly through autonomy (beta = 0.029), competence was the strongest predictor of performance (beta = 0.358) and relatedness had no independent effect. What the instruction contains therefore matters more than whether learners get prompting practice, and it matters most when the task is open-ended.

The variables surrounding AI literacy have since been mapped more broadly. A 2025 systematic review of AI literacy's correlates synthesized 31 studies across 14 countries (N = 12,071) and found the most consistent associations in the affective and behavioral band: AI self-efficacy, positive AI attitudes, motivation and digital competence all moved with AI literacy, while AI anxiety and negative attitudes moved against it. Demographic variables barely registered, with age and socio-economic status correlating weakly or not at all. The review's caution is about instruments: the same studies that produced strong correlations from self-assessment also showed far weaker ones when AI literacy was tested rather than self-rated.

Hu (2026) adds an ordered account of what AI literacy predicts rather than what predicts it. Surveying 450 university students in mainland China who already use ChatGPT, the study found AI literacy related directly to continued use (beta = 0.16, p = 0.002) and indirectly along a serial path through Trust and academic Self-Efficacy (indirect effect = 0.07, 95% CI [0.04, 0.10]), with the literacy to trust link the largest association in the model (beta = 0.50). AI anxiety moderated that first link (interaction beta = -0.25, simple slopes falling from 0.76 to 0.25 across the anxiety range), so the chain thinned for more anxious students and literacy instruction alone should not be assumed to reach them. The design is cross-sectional and the sample is restricted to existing users, so the ordering is a modeling assumption rather than an observed sequence.

A core applied aim of AI literacy is Reducing AI Misuse: teaching students to use AI ethically and productively rather than substituting it for their own cognitive work. Where AI literacy builds the capacity to evaluate and use AI critically, reducing misuse is the behavioral and structural payoff — combining guardrailed tool design, assessment redesign, and educative levers such as scaffolded think-first/AI-second sequences and prompting practice with deliberate Feedback. The two concepts are mutually reinforcing: AI literacy supplies the critical dispositions that make misuse-reduction interventions durable, while misuse-reduction evidence (e.g. the performance–learning gap) motivates why literacy must go beyond operational skill to critical judgment.

AI literacy also needs developmentally appropriate forms for the youngest learners. AI-Play translates AI literacy competencies into play-based, unplugged activities for Pre-K–K2 learners — organized around AI Body (AI as a system built from parts), AI Food (AI learns from examples), AI Brain (AI improves through patterns and feedback), and a Pre/Post-AI ethical lens — addressing a persistent lack of developmentally grounded AI literacy guidance for early childhood and making AI literacy accessible to non-technical educators and families. Complementing this, Vahedian Movahed & Martin (2025) found children aged 6–14 broadly trusted an age-tailored chatbot as an information source and most modeled it as "a smart computer program that learns" (nascent Machine Learning understanding), yet showed gaps in critical engagement and digital-safety awareness — evidence that young learners' trust can outpace their critical evaluation skills, making explicit Trust Calibration and safety instruction a necessary part of AI literacy for children.

Measurement is keeping pace with that developmental turn at the younger end. Thianwan and Srikoon's 2025 validation of an AI literacy self-assessment questionnaire for upper primary students developed a 15-item instrument for Grades 4 to 6, organized around Learning About AI, Learning About How AI Works, and Learning for Life with AI, and confirmed a three-factor structure in two samples (n = 335 and n = 579) with an overall Cronbach's alpha of .934. The authors frame it as a formative diagnostic of perceived rather than demonstrated literacy, and note that children's self-assessment accuracy is constrained by metacognitive ability that is still developing.

Project-based learning as a delivery mechanism. Zhu & Kong (2026) developed and validated an AI project-based learning (AI-PBLS) scale and, in a Hong Kong sample of 1,027 secondary and university students (446 with complete data), used structural equation modeling to show that empowerment in using AI for problem solving and AI ethical awareness jointly mediate the relationship between perceived Project-Based Learning and satisfaction with an AI literacy course. This positions PBL not merely as a delivery format but as a mechanism that builds learner confidence and ethical reasoning alongside competence, reinforcing the link between PBL and meaningful AI literacy development. It also supplies a validated measurement instrument for future measurement of AI literacy course experiences.

AI literacy as a core gap in Conversational AI frameworks. The umbrella review of conversational AI agents (Ganguly et al. 2025, 34 reviews) identifies limited AI literacy support as a major gap in CAI frameworks, and its ethical-use roadmap makes foundational assessment (including strengthening AI literacy) the first pillar alongside participatory design, ethical-use guidelines, and continuous evaluation of cognitive impact. It further finds that AI-literacy, training, and awareness rank among the most-emphasized ethical directions in the CAI literature.(Conversational AI agents in education: an umbrella review of current utilization, challenges, and future directions)

Effectiveness of AI literacy interventions. A three-level meta-analysis of 59 studies (172 effect sizes, 7,211 participants) estimates a large overall effect of AI literacy interventions (g = 0.837, p < .001) — but with a wide prediction interval [−0.292, 1.966], so effectiveness varies considerably across settings. Interventions in East Asia and Europe outperformed those in North America, and knowledge-focused interventions outperformed those targeting skills, attitudes, or ethics. The authors argue AI literacy education should therefore move beyond knowledge toward skills, practices, ethics, and attitudes, supported by integrated and reflective pedagogies (project- and problem-based, inquiry-based, experiential) and GenAI-supported tools — a shift that aligns with the participatory, producer-oriented forms of Computational Thinking described elsewhere in this knowledge base.

A second synthesis reaches a similar magnitude on a narrower base and sharpens the measurement problem. Yu, Kim, Chang and Huang (2026) pooled 57 effect sizes from 16 studies covering 3,837 K-12 students who were taught AI itself rather than taught with AI, and estimated Hedge's g = 0.892 (95% CI [0.548, 1.236], p < .001), robust to assumed pre-post correlations from 0 to 1 and stable after trim-and-fill at g = 0.952. Every one of the 57 effects pointed positive (range 0.16 to 4.93), and no moderator (publication source, year, school level) explained the substantial heterogeneity (I² = 94.68%). The authors' preferred explanation is measurement: the included studies called everything from AI knowledge and ethics to attitudes, motivation, self-efficacy and career interest "AI literacy," so the pooled figure describes a distribution of positive findings under one loose label rather than one uniform intervention. Read alongside Liu et al., the two estimates agree that AI literacy instruction works and agree on little else, which is itself the field's measurement signal.

Classroom evidence for deliberately brief designs is thinner, but it runs in the same direction and adds a behavior-level outcome that surveys cannot supply. Clerc et al. (2026) gave 116 French students in grades 8–9 a two-hour workshop on how LLMs work and fail, paired with practice in predicting whether a prompt would elicit a usable answer, evaluating the response, and revising or asking again. Two days later, during six LLM-supported science problems with a control group, trained students accepted underspecified prompts less often (51.5% vs. 66.7%, OR = 0.47), asked follow-up questions after a weak response far more often (59.2% vs. 27.9%, d = 0.80), and judged answer correctness more sensitively to prompt quality (interaction OR = 2.52), with a modest score difference (11.38 vs. 10.29 of 20, p = .040). The workshop never rehearsed the test tasks, so what transferred was a regulatory stance rather than task familiarity — and the authors are explicit that two days is not durability.

AI-supported critical media literacy in elementary school. Demir and Akar (2026) provide a concrete elementary-school demonstration of GenAI-supported literacy instruction: an 18-hour, 5E-model program for fourth-grade Turkish students in which ChatGPT and Grammarly were embedded phase-by-phase as pedagogical agents (ChatGPT for reflective questions and Q&A, Grammarly and Canva AI for content refinement, Padlet for peer feedback), aligned to the Turkish Language and Social Studies curricula. The AI-supported group showed large gains in media reading (+3.50), writing (+1.67), and total media literacy (+5.17, all p < .01) with between-group effect sizes of Cohen's d = 1.12–1.31, while the control group advanced only modestly. Qualitative analysis surfaced six domains of critical media literacy growth — digital self-protection and data privacy, purposeful and responsible media use, safe communication and boundary awareness, critical evaluation and misinformation awareness, online risk awareness, and media ethics/digital citizenship — evidence that developmentally appropriate, discipline-embedded GenAI use can build the critical-evaluation and ethical dimensions of AI literacy, not just operational skill.

Frameworks for structuring AI literacy. Several recent contributions offer structured progressions for building AI literacy. Marienko, Markova, and Semerikov (2026) propose a five-level framework (Awareness, Application, Evaluation, Creation, Ethics) integrated with three paradigms of AI in education (AI-directed, AI-supported, AI-empowered), developed through a mixed-methods study of Ukrainian secondary educators (national survey n = 2018; PD evaluation n = 1130). They found 84% of educators use AI but only 11% can identify specialized services beyond ChatGPT, and a professional-development intervention produced a 24% improvement in AI competence — evidence that targeted PD advances literacy beyond surface-level tool familiarity. The framework's grounding in Constructivism, connectivism, and TPACK connects to Technological Pedagogical Content Knowledge (TPACK) and Teacher AI Competency. Complementing this, Moore et al. (2026) used a two-year DBR process with a youth and AI-expert advisory board to design a science-integrated ML curriculum for high school youth, finding ML-knowledge gains in both cohorts (Cohort 2 M2−M1 = 0.175 vs Cohort 1 0.076) and greater gains among female and non-White participants — evidence that participatory, discipline-integrated design can advance both AI literacy and Equity.

The construct is contested precisely where educators are concerned. A review of 34 studies on AI literacy in teacher education found that the field lacks consensus on what AI literacy means for professional educators as distinct from general digital literacy, and UNESCO's AI Competency Framework for Teachers specifies 15 competencies across five dimensions while functioning as a global policy reference rather than implementation guidance. The faculty in the cross-institutional teacher preparation collaboratory treated AI literacy as an integrated dimension of professional preparation rather than a separate technology skill, so candidates were expected to audit, compare and justify AI output within disciplinary coursework rather than to demonstrate tool familiarity.

Initial teacher preparation is where the structure is thinnest. Pinto et al. (2026) reviewed 11 studies of AI in initial teacher training for pre-service primary mathematics teachers and found that nine were a single session or a handful of sessions embedded in existing courses, prompting was explicit training content in only some of them, and Ethics appeared in just three, none of which addressed transparency, algorithmic bias or accountability. The risks they report are literacy-level failures: pre-service teachers failed to notice conceptual errors produced by ChatGPT, accepted generated content uncritically, and delegated more problem solving to the tool as the tasks became harder. Their recommendation is to build AI competency across the whole training sequence, from a preparatory phase into the practicum, rather than in an isolated module.

Teachers' accounts of why AI education is taught add a purpose layer that competency frameworks leave implicit. Fagerlund et al. (2026) interviewed 13 Finnish teachers from preschool through grade 9, largely novices in AI education, and read their purpose statements through Biesta's three domains. Qualification, the acquisition of competencies, was the clearest and most concrete purpose, and it operated as a pathway into the other two: students were to grasp AI as a sociotechnical phenomenon and use its tools, with AI serving both as a target and as an aid to learning. Socialization appeared as cultivating enthusiasm, presenting AI as "cool," and framing students as future AI-savvy workers who must meet competency demands. Subjectification, the self-determined and personally meaningful engagement Biesta treats as central, was named as important but had no concrete instructional repertoire beyond discussing and contemplating, and teachers offered no worked examples for contesting AI's authority. The authors propose informed AI agency, a heuristic that places subjectification at the center while treating competencies as explanatory frameworks for self-determined choices. The design implication is that the "why" of AI education needs as much design attention as the "how": when subjectification rests on generic discussion, AI literacy instruction reverts to the qualification it was meant to sit beside.

Readiness itself is a contested target. Charles (2026) argues that teacher readiness for AI-rich learning should be read from the quality of the designs teachers produce, not from adoption, confidence or digital competence, and builds a framework around design mediation: observable decisions such as tool selection, task and prompt design, scaffolding, verification, transparency, assessment redesign, accessibility planning and human oversight. Two of its claims have direct professional-development consequences. First, Self-Efficacy is a mobilization resource rather than a quality indicator, because confidence without AI-specific pedagogical knowledge or ethical orientation can accelerate poor design rather than prevent it, which qualifies the usual reading of teacher confidence as a readiness signal. Second, policy clarity and institutional support are separate moderators, so material support can be high while expectations stay ambiguous and readiness still translates poorly. The framework also declines to equate non-use with unreadiness, since a reasoned decision to withhold AI can itself demonstrate design judgment. It is conceptual and unvalidated, so it supplies propositions and evidence sources rather than established effects.

AI-interaction literacy: the interactional dimension. Brunnström and Palmqvist (2026) propose a narrower, interactional competence — the ability to steer, evaluate, and learn from iterative interaction with GenAI — as a specific enactment of the applicational, evaluative, and integrational dimensions in broader frameworks such as the AI Literacy Heptagon. Their reflective demonstration shows what it consists of in practice: recognizing that a fluent answer is pitched above one's own schema, requesting simplification, narrowing scope, and redirecting the system toward focused practice. Two design consequences follow — the skill is unevenly distributed, so unguided use may advantage already-confident students and widen gaps (Equity), and it has to be taught explicitly rather than assumed, with teachers covering how to formulate productive prompts and when to stop using the tool (Self-Regulated Learning).

Critical AI literacy: beyond skills to power and resistance

A distinct strand of the knowledge base treats AI literacy not merely as skills or critical evaluation but as a critical and political practice that interrogates power, authority, and whose knowledge counts. This connects AI literacy to Critical Pedagogy and Equity:

These critical strands complement the operational and cognitive dimensions of AI literacy: where the latter ask "can the learner use and evaluate AI?", critical AI literacy asks "does the learner understand and challenge the power structures AI embodies?"

Whose AI skills count? The instructor–employer framing divide

A central open question in AI literacy is which skills matter and for whom. An Ithaka S+R study comparing how US instructors and employers prioritize the 26 skills of the HiBob AI Skills Framework found they agree on the importance of only one (setting realistic expectations for AI-augmented work). Instructors weight a critical, responsible-use orientation — recognizing AI's limits, transparency and attribution, human accountability, proactive output review — which aligns with academic values of attribution, review, and information literacy. Employers weight productivity-oriented skills — workflow evaluation and redesign, automation, and human–AI teaming — that reflect team-based workplace efficiency. Because only three of 26 skills are taught by half or more instructors, and those taught skew toward the critical-use categories, the report identifies a concrete AI skills gap: whole categories employers value are neither prioritized nor taught in college curricula.(AI skills for college graduates: Exploring how instructors and employers prioritize AI skills differently) This divide frames AI literacy as a contested construct — critical-use literacy for academic settings versus workflow-integration literacy for employment — a tension relevant to Framing AI Use for Students, Curriculum Design, and Workplace Learning.

Students' own ratings tell a related but differently ordered story. ElSayary and Ragab (2026) surveyed 380 university students, mostly STEM-enrolled frequent AI users, and found self-reported AI literacy carried most of the explanation of perceived employability readiness: four literacy dimensions added ΔR² = 0.554 (p < .001) on top of background controls worth only 0.054, and the full model reached 60.8% of the variance. Within the literacy set the evaluative dimension was strongest (Analyze & Evaluate, β = 0.286), followed by the affective one (Attitudes & Mindsets, β = 0.266), with hands-on use (β = 0.217) and foundational knowledge (β = 0.187) behind them; for perceived career readiness specifically, only Analyze & Evaluate (β = 0.216, p < .001) and Attitudes & Mindsets (β = 0.192, p = .004) were significantly associated. Read against the Ithaka divide, students' own sense of workplace readiness leans toward the critical-evaluation orientation instructors favor rather than the workflow-integration skills employers weight. The design is a single-session self-report with uniformly high means (4.17 to 4.57 of 5) and no employer or employment measure, so what it shows is that self-assessed AI literacy and self-assessed employability move together, with the evaluative and dispositional parts doing most of the work, not that these students are demonstrably workforce-ready.

Designing AI literacy interventions

The knowledge base's frameworks and empirical studies converge on a set of practical guidelines for educators and instructional designers building AI literacy interventions:

1. Use a structured competency framework as scaffolding, not a checklist. Mature frameworks give designers a shared vocabulary and a developmentally sequenced target. The SAIL framework organizes AI literacy into three domains (AI Concepts; Application and Technical Skills; AI Digital Citizenship) across four scaffolded levels — Understand and Explore → Apply and Integrate → Evaluate and Create → AI++. The AI Literacy Heptagon cross-cuts seven dimensions (technical knowledge, application, critical thinking, ethics, social impact, integration, legal/regulatory) with four Bloom-aligned proficiency levels, and stresses that emphasis must be adapted to disciplinary context — technical programs weight application, humanities programs weight ethical and social-impact reasoning.(The Scaffolded AI literacy (SAIL) framework: Results of a Delphi study for equitable AI literacy framework design in education)(The AI Literacy Heptagon: A Structured Approach to AI Literacy in Higher Education)

2. Engage learners across ICAP modes. Applying the ICAP Framework, effective instruction gives learners opportunities to engage at multiple cognitive levels — passive exposure (AI concept lectures), active manipulation (hands-on tool use), constructive generation (creating AI artifacts, self-explaining), and interactive dialogue (collaborative Problem Solving with peers and AI) — selecting the mode that fits the learning goal. A systematic review found successful collaborative AI-literacy interventions spanned all four ICAP modes.(Systematic Review of Collaborative Learning Activities for Promoting AI Literacy)

3. Assess demonstrated competence, not self-perception. Self-reported AI literacy diverges sharply from measured performance — teachers overestimate their AI skills by ~40%, and performance-based measures correlate with classroom AI integration far better than self-reports (r≈0.72 vs 0.31). Design interventions around performance-based assessment and calibration rather than confidence surveys, and use diagnostic profiles (overestimators vs. true novices) to target support.(How to Assess AI Literacy: Misalignment Between Self-Reported and Objective-Based Measures)

4. Build metacognitive and critical dispositions, not just operational skill. AI literacy is better understood as a metacognitive social practice than a skills checklist: because LLMs are probabilistic and opaque, learners must monitor and adjust their strategies, cultivate scientific skepticism, and interrogate how algorithms shape knowledge production — not merely learn to operate tools. Participatory co-design and experimental, project-based spaces (not one-off tool training) are where this awareness grows.(Metacognitive AI literacy: going beyond the AI skills gap agenda)

5. Embed literacy in the discipline and make it sustained. Movement toward higher stages of AI literacy (from uncritical use → informed use → critical evaluation → improvement) is most visible when experiences are sustained and discipline-embedded rather than delivered as standalone workshops. Design for repeated, contextual practice within authentic coursework.(Beyond Tool Adoption: A Practical Five-Stage Developmental Continuum for AI Literacy in Higher Education) A concrete model of discipline-embedded critical literacy is the task-based taxonomy of Dierickx et al. (2026) for journalism: it maps LLM-supported tasks across the four stages of the news workflow (newsgathering, sensemaking, editing, publication/distribution), each tied to a baseline prompt and a risk-and-mitigation strategy. By treating task definition and prompting as a situated form of professional judgment — not a neutral technical skill — it turns prompting itself into a vehicle for critical AI literacy (bias, hallucination, overreliance, and the enduring value of human editorial oversight), and its underlying logic transfers to other knowledge-intensive professions (law, medicine, public policy). Complementary to such discipline-specific taxonomies, the Dohn et al. (2026) taxonomy classifies GenAI learning activities along six dimensions, of which Epistemic Engagement (understanding / using / critiquing / constructing GenAI) directly operationalizes AI literacy as the depth of learners' cognitive relationship with the technology — from passive exposure to active critique and construction.

6. Treat equity and the digital divide as design constraints. AI literacy is a mechanism for addressing the three-level digital divide (access, skills, outcomes): closing the device gap is insufficient unless skills and critical use are built so benefits distribute fairly. Interventions should plan explicitly for learners who enter with less prior AI access, and incorporate cultural and governance perspectives rather than treating literacy as culture-neutral.(The Scaffolded AI literacy (SAIL) framework: Results of a Delphi study for equitable AI literacy framework design in education)(Digital Divide)

7. Pair literacy with misuse-reduction levers. Because AI misuse actively harms durable learning, literacy instruction should be coupled with the structural and educative levers documented under Reducing AI Misuse — guardrailed "hint-not-answer" tool design, assessment redesign, and scaffolded think-first/AI-second sequences with deliberate prompting practice.

8. Start from learners' actual entry point. Learners enter with distinct orientations — avoidance driven by fear, mistrust, or lack of access, versus uncritical reliance that masks misunderstanding. A diagnostic, stage-based approach (rather than a uniform curriculum) lets designers meet students where they are and move them toward critical, responsible engagement.(Beyond Tool Adoption: A Practical Five-Stage Developmental Continuum for AI Literacy in Higher Education)

9. Extend provision beyond the audiences formal education already reaches. The AI Literacies framework for public service media argues from a landscape review of more than 40 frameworks and 35 expert interviews that provision has tilted toward technical and functional skills, and that AI literacies overlap and should be connected to digital, media and information literacy rather than taught as a standalone subject. Its structure — six competency areas, five values and three progression levels (Understanding and Applying, Analysing and Evaluating, Synthesising and Specialising) with an assessment guide — is a usable template for providers outside formal schooling, and its equity argument is a design instruction: reaching young people who are digitally or otherwise marginalized takes deliberate partnerships, not universal publication.(Supporting AI Literacies for Young Adults Aged 14-19: A value-based, practical framework for public service media organisations)

Measuring AI literacy

A distinct research thread treats AI literacy not only as a target for instruction but as a construct to be measured. The knowledge base's assessment strand distinguishes self-reported from performance-based literacy: self-reports diverge sharply from demonstrated competence (teachers overestimate by ~40%), and performance-based measures predict classroom AI integration far better than confidence surveys (r≈0.72 vs 0.31). Intervention work reproduces that divergence at the level of behavior: in Clerc et al. (2026), neither GenAI attitudes nor a general metacognitive-awareness scale predicted students' regulation of LLM interaction or their final task scores (r = .01 and r = .04, both non-significant), while the behaviors the workshop changed — rejecting underspecified prompts, judging answer correctness, asking a follow-up — did track answer quality. Validated instruments are emerging to close this gap — the GLAT provides a psychometrically validated generative-AI literacy assessment, and diagnostic profiles (overestimators vs. true novices) let designers target support where it is needed. For design and research, this ties AI literacy to Educational Measurement and to Assessment broadly: a literacy framework is only as useful as the instruments used to track growth, and stage-based continua require reliable measurement to place learners along them.

Zhi, Yang and Huang (2026) build a graduate-specific model on Marzano's taxonomy: grounded-theory analysis of interviews with 14 professors condensed 329 raw labels into 96 concepts, 15 categories and five dimensions (cognitive foundation, operational skills, higher-order thinking, metacognitive reflection, ethical responsibility), then operationalized them as a 15-item Likert scale whose five factors emerged in exploratory factor analysis (83.14% of cumulative variance) and held in confirmatory factor analysis on a second subsample (CFI = 0.934, RMSEA = 0.083), with reliability from 0.796 to 0.842 across 308 valid questionnaires.

A related question is what the instruments can measure at all. Burriss et al. (2026) note that existing AI literacy scales and competency frameworks assume individually measurable performance and so structurally exclude collaborative, creative expression — their unit's evidence was Multimodal AI film artifacts, reflections, and civic discourse rather than a summative score, and the authors argue such evidence can complement rather than replace conventional measures. Broadening the construct may therefore require broadening the admissible evidence, not merely adding modality-rich items to existing scales.

A complementary strand measures how learners and teachers receive AI literacy materials rather than their literacy itself. Wang, Chuang and Wu (2026) had 794 students and 37 teachers rate two guidebook editions built to UNESCO's age threshold (ages 9-12 and 13-18) after roughly 30 minutes of guided classroom exposure. Acceptance held a four-factor structure (Performance Expectancy, Effort Expectancy, Perceived Playfulness, Behavioral Intention) with measurement invariance supported across the two student editions, younger learners scored higher on all four constructs, and perceived playfulness carried the largest association with intention in both cohorts. The authors are explicit about what this is not: perceived acceptance of an age-tiered resource is not AI literacy achievement, ethical reasoning, adoption or sustained use, and material-level acceptance should not be read as evidence that literacy improved.

Measurement also extends to the educators who mediate learners' engagement with AI. Most AI-literacy assessments target students or general users, leaving a gap in teacher education — a gap the Teachers' AI Literacy Scale (TAILS) addresses: grounded in the ED-AI framework with six dimensions (knowledge, evaluation, collaboration, contextualization, autonomy, and ethics), it was validated through exploratory and confirmatory factor analysis with preservice language teachers. Such instruments support measuring and developing the AI literacy of the educators who mediate learners' engagement with AI. On the student side, Nie et al. (2026) develop and validate a Generative AI Assessment Literacy Scale (GAA-LS) for higher-education students — an 18-item, five-factor instrument whose scores track Feedback engagement, Academic Integrity, and responsible AI use, with an indirect path to integrity running through feedback — evidence that assessment-specific AI literacy is measurable and ties to integrity behavior. What that teacher-facing instrument corpus actually contains has since been audited: Zainal, Mohd Matore and Maat's 2026 review of teacher AI literacy measurement tools appraised 33 instruments published between 2019 and 2025 and found that 31 (93.9%) were self-report scales of perceived confidence, only two tested knowledge objectively, and none used performance-based tasks, while fairness evidence was the weakest quality domain, with only five instruments reporting measurement invariance or differential item functioning.

The 2026 update of that measurement literature reorganizes it into four domains — knowledge and use, epistemic oversight, reliance calibration, and operational control of tool-using agents — and reports that no validated individual-level instrument in the corpus covers the full combination of scope, permissions, recovery, state isolation, independent review and evidence-based closure that agentic tool use demands. Its pooled subjective–objective correlation across three same-sample effects was r = .055, consistent with the divergence described above rather than with self-ratings as a usable proxy (Verí (2026)).

The measurement landscape is itself disordered, and now partly mappable. He, Zhang, Wang and Ji (2026) analyzed the AI literacy instrument corpus with an LLM-based coding pipeline and found the field measuring many constructs under one label — a jangle problem, with one construct carrying several names (Behavioral Commitment appears in eight candidate pairs, Intrinsic Motivation in five), alongside candidate jingle cases where instruments share a label but not their item content. Their semantic-similarity recovery of item and construct structure correlated only moderately with instruments' own reported reliabilities (r = 0.49 at item level against 0.45 and 0.35 at construct level), and they offer the method as screening rather than adjudication. The consequence for anyone selecting a measure is that convergent claims across studies are not safe to assume: two instruments with the same name may not operationalize the same construct.

What those instruments can support has now been audited, and the audit is sobering. Jin, Gašević, Martinez-Maldonado and Yan (2026) screened 8,056 records down to 58 studies covering 47 unique instruments, appraised with COSMIN against a construct whose instrument-development window compressed into three years (two instruments before 2023, 28 in 2025 alone). The corpus remains 37-of-47 self-report, and the evidence clusters tightly around internal structure: structural validity was sufficient in 32 of 58 studies and internal consistency in 32 with none rated insufficient, but construct validity was reported by only 15 studies (12 sufficient), criterion validity by one, measurement invariance by five (all sufficient), and just 2 of 58 earned a sufficient overall content-validity rating — comprehensiveness was insufficient in 53 despite 43 studies reporting cognitive interviewing. The pattern is that measurement practice looks robust exactly where it is cheapest to demonstrate and thin where evidence must come from outside the instrument, so a clean factor structure and a high coefficient (which can reflect item redundancy as much as construct representation) should not be read as demonstrated capability. The authors' prescription is consolidation rather than proliferation: refine existing scales, test invariance whenever one crosses a language or an educational stage, and add criterion evidence by relating scores to performance tasks.

One of the exceptions is built to compare groups. Zhang et al. (2026) developed the Generative AI Literacy Scale (GAILS) — 43 items reduced to 34 through a five-expert Delphi review and a seven-person pilot, then validated in 341 North American adults — and tested scalar invariance, finding that it held across female and male respondents and across student and workforce groups. That is what licenses the mean comparisons the instrument reports — men scoring modestly higher on Adaptive Operational Skills, students higher on Adaptive Operational Skills and Responsible GenAI Literacy — as competence differences rather than differential item functioning, and it is what makes the scale usable across Higher Education and workplace settings rather than student-only. Its own evidence is not unblemished: fit was mixed (CFI = .967 and SRMR = .058 inside conventional cutoffs, RMSEA = .087 above them), total-scale α = .973 sits alongside one failed Fornell–Larcker discriminant comparison (Factor 1 √AVE = .828 against r = .873 with Factor 3), and, like almost everything in the corpus, it measures perceived competence with no behavioral criterion. A validated scale in this field is a starting point for the invariance and criterion work the review calls for, not a finished Benchmark.

Connections across the knowledge base

AI literacy intersects with AI Tutoring (understanding when and how AI tutors are effective), Teacher AI Competency (educator preparedness), Academic Integrity (knowing what constitutes appropriate AI use), and AI in Education broadly. It is both a prerequisite for effective AI use and an outcome of well-designed AI integration — students learn AI literacy BY using AI critically, not just by learning ABOUT AI.

AI literacy is double-edged for overreliance: Maizel et al. (2026) found the skill-based dimensions of AI literacy (using/understanding, detecting) were positively associated with reported AI dependency, while AI Self-Efficacy and academic confidence were negatively associated — so technical AI-literacy training, absent self-efficacy and Self-Regulated Learning scaffolds, can increase dependency. AI literacy here becomes an enabling capacity whose direction depends on complementary motivational resources.

  • Verification habits pay off only through regulation. In Davor, Larbi and Boateng (2026)'s survey of 533 university students in Ghana, AI verification literacy had no significant direct effect on critical thinking (beta = .076, p = .090) or technical Problem Solving (beta = .043, p = .385) and mattered only through metacognitive self-regulation (beta = .167, p < .001), a full mediation the authors call metacognitive activation. AI task scaffolding predicted both outcomes (beta = .185 and .170), while cognitive offloading tendency predicted them negatively (beta = -.240 and -.312) and depressed regulation as well (beta = -.294). Teaching students to check AI output is therefore not enough on its own: evaluative habits need planning, monitoring and reflection built into the task.

  • Critique of AI output as a literacy practice: Hosseini (2026) treats evaluating AI-generated errors as a core AI-literacy skill, using failure-mode analysis and iterative prompt refinement in a database design course. The study found students overestimated their AI abilities (self-reported literacy weakly, negatively correlated with objective competency), and that critique-based learning strengthened calibration.

  • Socialist humanist AI literacy (2026): A literature review critiques compliance-oriented AI literacy and proposes a socialist-humanist framing of asynchronous AI literacy and fair use in higher education, linking the historical digital divide to modern AI literacy and calling for approaches that serve human flourishing and equity rather than mechanical policy compliance (From Mechanical Compliance to Human Flourishing: A Socialist Humanist Approach to Asynchronous AI Literacy and Fair Use in Higher Education).

  • Cheap, light-touch warnings blunt AI persuasion — without costing general trust. Orchinik and Rand (2026) preregistered two experiments (total N = 3,208 US adults) in which participants conversed with an Large Language Models (LLMs) instructed to shift their views on political topics. A single brief warning — that models can be prompted to persuade and may present information selectively — cut belief change by roughly one half (-48.1%, 95% CI [-59.5%, -36.8]) relative to control, and adding specific persuasion-technique warnings produced no further benefit. The property that matters for teaching is on the other side of the effect: general trust in Generative AI did not fall, so the intervention builds calibrated trust rather than blanket skepticism. It is a one-paragraph, no-facilitation intervention, which is a rare cost profile for an AI literacy design.

  • Learner governance of AI matters more than the design of the tool. Azimi (2026) randomized 33 students in a master's data-analytics course between a scaffolded AI Study Coach embedded in the notebooks (n = 16) and unrestricted use of any Generative AI tools they chose (n = 17) for seven weeks. Assignment performance and concept-inventory gains were indistinguishable; the Coach condition reported higher confidence instead. What separated students was AI literacy in practice: those who had formulated their own rules for when to use AI scored higher in both conditions, and the students with the deepest model understanding — every one of them self-taught — prompted most deliberately. The design implication runs against the control reflex: teach the self-regulatory and model-understanding components of AI literacy rather than constrain tools.

Connected Concepts

Connected Articles

Connected FAQs

Connected Resources

  • EduGems
    'A growing library of pre-made Google Gemini 'Gems' — shareable, pre-written prompts for planning, feedback, forms, images and AI-use reflection.
  • Pressing Prompts
    Openly accessible higher-education activities and resources for pressing critical questions about AI, organised into three thematic clusters.
  • mglearn Classroom Resources
    A large set of free, browser-based classroom activity suites and print-ready materials from TCEA, including Gen AI literacy and digital citizenship modules.
  • Education Agent Skills
    A library of 165 evidence-grounded agent skills covering pedagogy, learning science, curriculum, and assessment, packaged for Claude, Codex, and Hermes.

Embed this page

Copy the code below to embed a chromeless version of this page in a learning management system or other website. The embedded view hides the site header, navigation, and footer.