On this page

Language Learning — the study of how AI supports second language (L2) acquisition, writing development, and linguistic diversity in educational settings. AI in education research in this knowledge base spans AI interlocutors for spoken dialogue, automated writing evaluation for L2 learners, reading support, and concerns about language bias in AI scoring systems.

Questions to Consider

  • Language is inherently interactive, which makes it well-suited to conversational AI — but AI's linguistic capabilities also raise risks of bias against non-native patterns. Where have you seen this tension between opportunity and risk play out?
  • One study found AI scoring systematically underestimates linguistically weaker students, while another proposed comparing students to their own prior work rather than native-speaker norms. How does the 'reference point' for evaluation change whether AI feedback helps or penalizes a learner?
  • AI interlocutors can extend communicative practice at scale, but the page warns they should pair with human interaction so fluency transfers to real conversation. What might you gain from practicing with an AI that a human partner can't give — and what would you lose?
  • Teacher support — not just the AI tool — was shown to drive engagement in AI-assisted language learning through students' achievement goals. How does the social and pedagogical context shape whether learners keep engaging with an AI practice tool?
  • A meta-analysis found small-to-moderate, level-dependent gains from emerging tech, with productive skills (speaking, writing) gaining more than receptive ones. Why might speaking and writing benefit more than listening and reading from AI tools?
  • If AI privileges standard English and can penalize non-native or diverse language patterns, how should language instructors design evaluation and feedback so AI supports linguistic diversity rather than erasing it?

Introduction

Language learning has emerged as a significant AI in education domain because language is inherently interactive — making it well-suited to conversational AI — and because AI's linguistic capabilities raise both opportunities (personalized language practice at scale) and risks (systematic bias against non-native language patterns). The articles in this knowledge base explore both sides of this equation. Where the target language is English specifically — especially English for Academic Purposes (EAP) and EFL/ESL/L2 English teaching — see the dedicated English Education (EAP / EFL / ESL) concept page, which distinguishes English-specific research from general L2 acquisition and general writing.

AI as language tutor and interlocutor is the most developed theme. What Changes When the Interlocutor Is an AI? examines interactional fluency and linguistic uptake when L2 learners converse with AI versus humans. TACT provides pedagogically adaptive ESL tutoring. Children's English Reading Story Generation via Supervised Fine-Tuning of Compact LLMs with Controllable Difficulty and Safety explores AI-generated stories for children's reading development. These connect to Intelligent Tutoring and Generative AI. Yang, Weng and Yang (2026) designed two Large Language Models (LLMs)-based agents — a conventional AI English teacher and one using the 5E framework (engage, explore, explain, elaborate, evaluate) for inquiry-based grammar learning. Across 37 ESL students in a randomized comparison, high-performing students responded positively to the AI teacher while low-performing students showed mixed attitudes, and the conditions differed in intrinsic motivation, cognitive change, and performance — indicating that LLM-agent design should be matched to learner proficiency.

AI in language assessment is emerging as LLMs support item generation and evaluation. Aryadoust and Wong (2026) compared prompt engineering against fine-tuning for automatic item generation in L2 listening assessment: iterative prompt refinement improved item quality but plateaued, while fine-tuning GPT-4.1 on the optimized prompt (holding prompt design constant) yielded further gains — a template for when assessment developers should invest in model adaptation over prompt iteration.

Automated writing evaluation for L2 learners evaluates AI's ability to assess non-native writing. Bannò et al. proposed a self-referential approach comparing student writing to their own prior work rather than native-speaker norms. Feser & Tschisgale found AI scoring systematically underestimates linguistically weak students — a finding that connects to Assessment Validity and Bias Mitigation concerns. Generative AI and linguistic diversity in academic writing and publishing: Perspectives from World Englishes explores how AI affects linguistic diversity in academic contexts.

Accessibility for language learners connects to Inclusive Learning: DysLexLens analyzed how dyslexic learners use AI for literacy support, and AI tools in Arab University English classrooms: Looking back and forward explored AI tools in Arabic-English classroom contexts. These studies connect language learning to Equity and Special Education.

Motivational mechanisms in AI-assisted language learning examine why learners engage with AI for language practice. Wang & Wang (2026) used goal-setting theory with 758 Chinese university English learners to show that teacher support enhances engagement in AI-assisted learning through students' mastery-approach and performance-approach goals (not avoidance goals) — evidence that the pedagogical and social context, not just the AI tool, determines whether learners stay engaged with AI-assisted language practice. This connects language learning to Motivation and Student Engagement.

GenAI-supported writing at the primary level. Lu et al. (2026) ran a nine-week opinion-writing program with 301 Grade 5 and 6 learners in Eastern China, with eight intact classes randomly assigned to the program or to conventional instruction. The program raised learners' ideal L2 writing self (adjusted mean difference 0.20) and academic buoyancy (0.17), and lifted behavioral and emotional engagement, but it did not move growth mindset, cognitive or metacognitive engagement, or rubric-scored organization — among writing dimensions only language use improved. Two features of the design matter for language teachers: prompting was taught explicitly, through a categorized bank of prompts tied to specific writing goals, and GenAI feedback was used alongside comparison with teacher feedback and repeated revision. The authorship gains learners reported rested on that instructional structure rather than on the tool by itself, and the authors name reduced self-monitoring and shortcut-oriented strategies as the standing risks.

Implications for language instructors

  • Emerging technologies yield small-to-moderate, level-dependent gains. A meta-analysis of 33 TEFL studies (N = 3,181) finds an overall effect of Hedges' g = 0.38 that rises with educational level (primary 0.29, secondary 0.35, tertiary 0.44), with VR/AR yielding the largest effects and productive skills (speaking, writing) gaining more than receptive skills — supporting the use of emerging tech, especially at tertiary level, while keeping expectations realistic.

  • Use AI to extend communicative practice, not replace it. AI interlocutors and adaptive ESL tutors expand interactional practice at scale — pair them with human interaction so fluency and uptake transfer to real conversation.

  • Prioritize feedback quality over quantity in ASR-supported speaking. Chen et al. (2026) find that accurate error correction and structured reflection tasks improve Feedback internalization and reflective behavior in college English speaking, while frequent ASR use and recognition accuracy boost motivation or reflection only partially — technical precision alone does not drive deeper cognitive engagement, and language proficiency moderates the gains (stronger learners internalize feedback more effectively). This argues for pedagogically sound feedback (e.g., articulatory explanations over simple error flags), scaffolded reflection, and proficiency-differentiated support.

  • Support learners' psychological adaptation to AI-assisted study. Wu (2026) tracks learners of Japanese over a semester and finds they sort into maladaptive, moderate, and positive adaptation profiles driven by the balance of technostress and resilience, with most learners gradually shifting toward positive adaptation and reporting higher Self-Efficacy and lower burnout — a signal to design AI-mediated language practice that manages technological strain, not just tool access.

  • Be alert to scoring and feedback bias against learners. AI scoring can penalize non-native patterns; linguistic-diversity research warns AI privileges standard English — use self-referential or human-moderated evaluation.

  • Support the full spectrum of learners. Dyslexia and accessibility studies and culturally responsive design (Arab-English contexts) show AI must be adapted to diverse learner needs, not assumed universal.

  • Educators value GenAI for preparatory work, not live classroom use. A PRISMA systematic review of 23 studies (Li et al. 2026) finds language educators most value GenAI for behind-the-scenes preparation — lesson planning, materials creation, and writing support/feedback — yet remain hesitant about direct, classroom-facing implementation, reflecting a theory–practice gap between approving AI in principle and using it live. Adoption is shaped by professional-identity, pedagogical, technical, institutional, and Academic Integrity factors, with educators falling on a spectrum from non-adoption to comprehensive integration; attitudes tend to evolve from initial insecurity toward confident, selective use with exposure.

  • Prepare language teachers' AI literacy. Systematic reviews find AI literacy among language teachers is a key gap — invest in teacher professional development alongside tool adoption. As AI reshapes language education, AI literacy is also crucial for teachers to engage critically with the technology: the Teachers' AI Literacy Scale (TAILS) was developed for language teacher education, operationalizing the six-dimension ED-AI framework (knowledge, evaluation, collaboration, contextualization, autonomy, Ethics) and validated with preservice English language teachers.

  • Four interaction profiles in a high-pressure bilingual task. Kuang, Li and Weng (2026) used eye-tracking, pen-recording and voice-recording with 22 interpreting trainees to show that students divide attention between AI output and their own note-taking in four distinct ways — Intensive Engagers, Fast Scanners, Traditionalists and Frequent Switchers — and that 58.3% of stage-level observations changed profile between the comprehension and production stages of the same task. Only comprehension-stage patterns predicted product quality, and the AI-heaviest cluster scored lowest on fluency of delivery and target language quality, which makes the case for teaching learners to describe and reflect on their own strategy rather than prescribing one way of working with the tool.

Connected Concepts

Connected Articles

Embed this page

Copy the code below to embed a chromeless version of this page in a learning management system or other website. The embedded view hides the site header, navigation, and footer.