On this page

Synthesis: Yu, Kim, Chang, and Huang meta-analyze 16 empirical studies that taught artificial intelligence itself, rather than using AI to support instruction, to K-12 learners. Searching 8 databases from January 2014 to March 2024 and extending the search by snowballing through October 2025, they identified 57 effect sizes covering 3,837 students in studies published between 2021 and 2025. The pooled effect of AI education classes on AI Literacy was Hedge's g=0.892 (95% CI: [0.548, 1.236], p<0.001), a large effect by Cohen's criteria, and all 57 individual effect sizes were positive, ranging from 0.16 to 4.93. Heterogeneity was substantial (I2=94.68%), yet moderators by publication source, publication year, and school level showed no significant differences, and trim-and-fill adjustment left the effect positive and significant (g=0.952). The authors read the consistency as support for teaching AI concepts to younger learners, and the unexplained variability as evidence that the field has not agreed on how to measure AI literacy.

Key Findings

  • Across 16 studies and 57 effect sizes covering 3,837 K-12 students, AI education classes raised AI literacy with a pooled Hedge's g=0.892 (95% CI: [0.548, 1.236]), a large effect.
  • All 57 extracted effect sizes pointed in a positive direction, ranging from 0.16 to 4.93, so the overall result does not rest on a subset of null findings.
  • Heterogeneity was substantial (I2=94.68%; Q(56)=752.67, p<0.001), so most observed variance reflected true differences among studies rather than sampling error.
  • Publication characteristics explained little: journal articles showed β=0.8047 (p=0.0008), other publication types did not differ significantly, and year was unrelated to effect size (b=-0.0381, p=0.840).
  • Robustness held: the pooled effect and standard error stayed at 0.892 and 0.160 across assumed pre-post correlations from 0 to 1, and trim-and-fill gave g=0.952 (95% CI: [0.809, 1.096]).
  • AI knowledge was the most common outcome (10 of 16 articles, 62.5%), while AI career interest appeared in 3 articles (18.8%), showing how differently the construct was operationalized.

How the synthesis was assembled

The study draws a deliberate line between AI education, which teaches AI concepts, principles, ethics, and development so that learners build AI literacy, and AI in education (AIED), which uses AI tools such as automated grading or personalized feedback to support teaching. Most prior meta-analyses synthesized AIED outcomes such as achievement and engagement, leaving evidence on teaching AI itself thin. The authors kept only pre-post or experimental quantitative studies with K-12 samples, published in English, and reporting a concept similar to AI literacy, including Computational Thinking, AI ethics, Self-Efficacy, and career interest. Two of the 16 studies used between-subject designs and the rest pre-post designs, so all standardized mean differences were converted to Hedge's g, and a robust variance estimation model handled dependence among multiple outcomes from the same study. The set comprised 8 journal articles, 2 dissertations, 5 conference proceedings, and 1 book chapter, pooled in R with a random-effects model.

A large and consistently positive effect

The pooled estimate of g=0.892, an average learning gain across 57 outcomes, is the paper's headline claim, and it held under several checks. Varying the assumed correlation between pre- and post-test scores from 0 to 1 left the pooled effect at 0.892 with a standard error of 0.160. Publication-bias screening was mixed: the funnel plot looked mostly symmetrical, but Egger's test indicated significant asymmetry (z=5.49, p<0.001), with a limit estimate whose 95% CI of [-0.061, 0.506] included zero. The authors attribute that to a few outlying effect sizes rather than systematic bias, and trim-and-fill supported them, returning an adjusted effect of g=0.952 (SE=0.073, 95% CI: [0.809, 1.096]). Substantial heterogeneity remained (I2=94.68%; Q(56)=752.67, p<0.001), so the summary describes a distribution of positive findings rather than one uniform intervention.

Moderators and the measurement problem

Three moderator analyses came back null, and the authors treat that nullity as informative. Publication source did not differentiate studies, publication year showed no trend (b=-0.0381, p=0.840), and school level produced no significant contrast (difference -0.1748, p=0.5071). They caution that small subgroups limit statistical power. Their preferred explanation for the remaining heterogeneity is measurement: the 16 studies operationalized AI literacy as everything from AI knowledge and ethics to attitudes, Motivation, Self-Efficacy, and career interest, mixing cognitive and affective outcomes. They connect this to Educational Measurement: the field lacks consensus on AI literacy's core dimensions and validated instruments, and short-term classes rarely capture gradual shifts in attitudes. This unresolved construct openness is both why so many studies qualified and why their results resist comparison, a recurring limitation in AIED research.

What this means for practice

  • K-12 AI education can be justified empirically: the pooled large effect held across school levels, publication types, and years, giving policymakers a defensible rationale for curriculum integration.
  • Design AI education around explicit AI literacy goals rather than tool practice, since the synthesized interventions taught concepts, principles, ethics, and applications.
  • Pair rollout with professional development for pre- and in-service teachers, which the authors name as a precondition for implementation.
  • Read single-study effect sizes cautiously, because the range from 0.16 to 4.93 partly tracks differences in what each study called AI literacy.

Limitations

  • Only 16 studies were included, fewer than traditional meta-analysis requires; the authors note this reduced the statistical power of the moderator analyses.
  • Heterogeneity was very high (I2=94.68%) and was not explained by the moderators tested, so the pooled effect may vary by intervention design, context, and measurement.
  • Cognitive and affective outcomes were combined under a broad definition of AI literacy, pooling measures with different properties into one estimate.
  • Two studies reported no participant breakdown by school level, and detailed grade levels were not analyzed.

Citation

Yu, Wonjin; Kim, Nari; Chang, Ammi; Huang, Wanju. (2026). The Effects of K-12 Artificial Intelligence Education in Enhancing AI Literacy: A Meta-Analysis. Journal of Computer Assisted Learning, 42, e70308. https://doi.org/10.1002/jcal.70308

Embed this page

Copy the code below to embed a chromeless version of this page in a learning management system or other website. The embedded view hides the site header, navigation, and footer.