On this page

Synthesis: This meta-analysis pools 40 quasi-experimental and experimental studies published between 2014 and 2025, covering 3367 participants, to estimate the effect of digital and AI technologies on foreign language Language Learning. Using Hedges' g and a random-effects model, the pooled effect is 0.962 (95% CI 0.765 to 1.159), a large positive effect. The authors organize eight moderators with the Technology-Organization-Environment framework. Only three show significant between-group differences: research design (QB = 9.635, p < 0.01), intervention duration (QB = 15.179, p < 0.01) and the number of technologies used (QB = 19.132, p < 0.001). Language skill type, educational level, sample size, learning setting and technology type show no significant differences. Heterogeneity is high (Q = 263.122, I2 = 85.178), and publication-bias checks support the pooled figure (Orwin's fail-safe N = 5853; Egger's test p = 0.84465). The authors conclude that effectiveness cannot be attributed to a technology label or a skill category, and that longer or more tool-heavy interventions are not automatically better.

Key Findings

  • Across 40 studies and 3367 participants, digital and AI technologies produced an overall pooled effect of 0.962 (95% CI 0.765 to 1.159, Z = 9.586, p = 0.000), which the authors classify as large.
  • Heterogeneity was high (Q = 263.122, I2 = 85.178, p = 0.000), above the 75% threshold the authors set for substantial variation, so the random-effects estimate is the reported one.
  • Research design moderated outcomes (QB = 9.635, p < 0.01): quasi-experimental studies reported g = 1.019 while true experiments reported g = 0.474.
  • Intervention duration moderated outcomes (QB = 15.179, p < 0.01) but not linearly: 5 to 8 weeks gave g = 1.112, 9 to 24 weeks g = 0.847, and under 2 weeks g = 0.163.
  • The number of technologies used moderated outcomes (QB = 19.132, p < 0.001), ranging from g = 0.882 for one tool to g = 2.205 for four, though only one study used four or five tools.
  • Writing showed the largest skill subgroup effect (g = 1.198) and reading the smallest (g = 0.747), yet skill type did not moderate effects (QB = 2.853).

How the corpus was assembled

The review searched Web of Science, Scopus, ERIC, the ProQuest Education Database and IEEE Xplore, then screened with PRISMA 2020. From 2294 retrieved records, 40 studies qualified: journal articles with a quasi-experimental or true experimental design, a control group taught by traditional methods, a pretest, and enough statistics to compute an effect size. Together they cover 3367 participants, each contributing one Hedges' g value. The corpus is skewed: higher education supplied 72.5% of studies, quasi-experiments 87.5%, classroom settings 75%, Technologies rather than general digital tools 80%, and single-technology designs 67.5%. Speaking and writing dominate the skill categories; listening and reading are scarce.

The overall effect

Under a random-effects Meta-Analysis and Systematic Review, the pooled effect of digital and AI technologies on foreign language skills was 0.962 (95% CI 0.765 to 1.159, Z = 9.586, p = 0.000), a large effect. Heterogeneity was substantial (Q = 263.122, I2 = 85.178, p = 0.000), so the random-effects figure is the one the authors interpret. Because the between-group test for skill type was not significant (QB = 2.853), this ordering cannot be read as a reliable ranking of skills.

Which moderators held, and which did not

Three of the eight moderators produced significant between-group differences. Research design came first: quasi-experiments reported g = 1.019 against g = 0.474 for true experiments. Intervention duration came second, and its four bands did not rise monotonically (g = 0.163, 0.999, 1.112 and 0.847). The number of technologies used came third, with the largest subgroup values appearing at three tools (g = 1.735) and four tools (g = 2.205); only three studies used three tools and only one study each used four or five. The other five moderators were not significant: language skill type (QB = 2.853), educational level (QB = 0.143), sample size (QB = 0.067), learning setting (QB = 1.840) and technology type (QB = 0.049). K-12 (g = 1.030) and higher education (g = 0.939) both improved, blended learning showed the largest setting effect (g = 1.165), and AI tools such as Conversational AI chatbots (g = 0.951) could not be separated from general digital tools (g = 1.009).

Robustness and interpretation

The funnel plot was asymmetric, and trim-and-fill estimated three missing studies; correcting for them raised the pooled effect from 0.962 to 1.033 (95% CI 0.835 to 1.231). Orwin's fail-safe N was 5853 and Egger's regression test was non-significant (t = 0.19730, p = 0.84465), so the authors treat the estimate as robust. They read the moderators through the Technology-Organization-Environment framework with cognitive load theory, social constructivism and Self-Regulated Learning theory, and caution that large effects in quasi-experiments, small samples and thin subgroups may be inflated. The closing discussion also raises the governance side of AI-supported language learning: Privacy concerns around learner data, and the risk that uneven access widens existing inequality.

What this means for practice

  • Select tools by instructional objective, learner proficiency and task type, since AI and general digital tools did not differ significantly (QB = 0.049).
  • Do not assume that longer interventions or more tools are better; the duration and tool-count patterns were non-linear, and several subgroups rest on one or two studies.
  • Read quasi-experimental effect estimates cautiously, as they ran more than twice the size of true-experiment estimates (g = 1.019 vs 0.474).
  • Pair adoption with data protection for learner records and attention to unequal access, the two risks the authors flag for technology-supported language learning.

Limitations

  • Some primary studies had small samples, notably in multi-technology and remote-learning contexts, which limits the generalizability of the results.
  • Several studies addressed only one educational level or one language skill, constraining the scope of the conclusions.
  • Restricting the search to 2014 to 2025 leaves long-term effects unexamined, and with I2 = 85.178 several moderator patterns rest on very few studies.

Citation

Chen, Yinong; Wei, Lina. (2026). The Impact of Digital and Artificial Intelligence Technologies on the Improvement of Foreign Language Listening, Speaking, Reading and Writing Skills: A Meta-Analysis. Journal of Computer Assisted Learning, 42, e70325. https://doi.org/10.1002/jcal.70325

Embed this page

Copy the code below to embed a chromeless version of this page in a learning management system or other website. The embedded view hides the site header, navigation, and footer.