On this page

Speech and voice technologies — the family of AI systems in which the spoken channel carries the interaction: automatic speech recognition (ASR) that transcribes, analyses, and grades what a learner says; text-to-speech (TTS) that generates narrated or dialogue-form instruction; voice-first agents that hold real-time spoken conversation with learners; and automated capture and scoring of oral performance. Where Conversational AI research is mostly about text chatbots and Language Learning covers second-language pedagogy broadly, this page follows the audio itself — what a speech interface changes about learning, what it costs, and who it includes or excludes when the interface is a voice rather than a screen.

Questions to Consider

  • ASR feedback helps stronger learners more than weaker ones: in one survey of 325 Chinese undergraduates, the payoff of reflective behavior and motivation was weak and non-significant at low proficiency and grew markedly stronger at average and high levels. If a tool raises the average while widening the gap, is it a success?
  • TTS narration matched an instructor's own voice on comprehension, concentration, and overall evaluation — but dialogue-format TTS beat single-speaker TTS on comprehension and cognitive engagement while sounding less natural. Which would you choose for a first introduction to a concept, and which for a dense procedural walkthrough?
  • Two findings run against intuition: students who read the AI's output most heavily during interpreting tasks had the weakest delivery fluency, and a learner who answered in two words was marked correct while a fuller spoken answer was flagged. How much does a voice interface measure verbal fluency rather than understanding?
  • A voice agent with a British accent was treated as a tool, while agents with Indian and African American accents were anthropomorphized and treated as peers in K-12 group work. If accent changes how a learner relates to an agent, who should decide which voice an educational product ships with?
  • Voice-first and offline systems exist mainly to reach learners that screen-based edtech excluded — blind children, deaf and hard-of-hearing learners, learners in workshops with no connectivity. What would it take for accessibility to be the starting assumption of a speech product rather than a later fix?
  • Nearly every claim on this page comes from a short, single-site, self-report study with a fixed presentation order. Which of these findings would you trust enough to change your assessment or lesson design before a replication exists?

Introduction

Speech is the oldest educational interface and the newest one for AI. Systems that listen and speak now sit in language classrooms, vocational workshops, and high-school inquiry lessons. They are attractive for the usual reason — speech is fast, works hands-free and eyes-free, and is the medium in which pronunciation, fluency, and oral reasoning live — and for a newer one: a time-limited spoken answer is hard to outsource to a text generator, which makes the spoken channel a candidate answer to Academic Integrity pressure on written assessment.

The research here splits into four clusters: ASR pronunciation feedback, TTS-generated instruction, voice-first and spoken-dialogue partners, and AI-supported oral Assessment. Running through all four is a question about equity — speech interfaces remove one barrier (sight, reading speed, typed literacy) while introducing others (accent, recognition accuracy, hardware, connectivity).

ASR and Pronunciation Feedback

The most consistent message from the ASR work is that the tool is not the treatment — the feedback design is. Chen et al. (2026) surveyed 325 undergraduates at a Chinese teacher-training university and modeled accuracy, usage frequency, feedback quality, and reflection-task design against reflective behavior and intrinsic Motivation, then against speaking improvement. Accurate error correction and structured reflection tasks drove both Feedback internalization and reflection. More frequent use boosted reflection but had no independent effect on motivation. Recognition accuracy raised motivation, plausibly by building trust in the tool, but did not by itself trigger deeper processing — a learner can register a flagged error without analyzing its cause. Reflection was the stronger predictor of speaking gains, and proficiency moderated both pathways: weak and non-significant at low proficiency, markedly stronger at average and high levels. The advice that follows is to explain errors articulatorially rather than flag them, ramp complexity with readiness, and frame correction supportively.

A second strand asks whether AI pronunciation feedback changes learners' disposition to speak. Lu et al. (2026) surveyed 1,701 Chinese university EFL learners and used covariance-based structural equation modeling with bias-corrected bootstrapping to test whether perceptions of Generative AI pronunciation feedback related to willingness to communicate, with pronunciation Self-Efficacy as mediator. The association was positive, and self-efficacy partially mediated it: the indirect path accounted for 69.9% of the total effect while a direct effect remained. All constructs were self-reported at one time point, so the study describes a mechanism learners perceive rather than one that was manipulated.

Automated pronunciation feedback can also be built without error labels. Kawamura (2026) describes Profy, which learns what good and poor pronunciation look like from largely unannotated speech via self-supervised learning, then shows learners which waveform regions drove the judgment and how far their acoustics sit from native-speaker distributions. With 10 Japanese learners of English rated by five American listeners, intelligibility improved, and unlike an elicited-imitation baseline the pre- and post-practice confidence intervals did not overlap — a small sample, but one that gives where and how much rather than a binary verdict.

TTS and Voice-Generated Instruction

Generated narration turns out to be a viable substitute for a recorded teacher, and the interesting variation lies inside the format rather than between human and machine. Kumoi et al. (2026) built a three-stage, human-in-the-loop pipeline in which an Large Language Models (LLMs) drafts slides and TTS-optimized narration while the educator fact-checks at each stage, then ran a within-subject study with 245 first-year students at a Japanese prefectural high school across instructor-voiced video, single-speaker TTS, and Expert-×-Novice dialogue TTS. Comprehension, concentration, and overall evaluation did not differ significantly across the three (p > .14, r < 0.11), and equivalence testing kept all three within a ±0.5 margin: synthetic narration did not degrade the experience. Dialogue narration was better on comprehension (p = .006, r = .130), on reported deepened thinking (p = .019, r = .121), and on being able to explain the content to a friend (p = .004, r = .152), and 66.9% preferred it as most enjoyable. Two caveats: dialogue audio was rated significantly less natural (r = −.238), consistent with extra load from multiple speakers, and the dialogue group arrived with less prior knowledge (66.0% reporting knowing "nothing at all" versus 34.1%), which makes its advantage a conservative estimate.

The follow-up argues that format should be matched to the learner, not the content. Watanabe et al. (2026) gave 222 first-year high-school students teacher–student (TS), student–student (SS), and teacher–teacher (TT) dialogue lessons generated by an LLM and voiced through a TTS API, fitting linear mixed-effects models through an aptitude-treatment interaction lens. TT raised Motivation more than TS, but the effect depended on learning style: the interaction with the Concrete Experience factor was positive (b = 0.162, p < .001) while the interaction with the reflection-and-conceptualization factor was negative (b = −0.238, p = .002). TT also drew a significantly lower overall evaluation than TS (b = −0.126, p = .005), which the authors attribute to intrinsic cognitive load from dense expert-to-expert exchange; free responses name "stiffness of speech" in TT and "unnatural manner of speaking" in SS. Both TTS studies confound format with content and fixed viewing order, so their effect sizes are preliminary.

Voice-First Companions and Spoken Dialogue Partners

What changes when the interlocutor is a machine? Scheinberg et al. (2026) had 78 university learners of German across four sites complete a counterbalanced spot-the-difference task with both a human peer and a real-time AI partner, then analyzed diarized ASR transcripts. Human dialogue was faster and better balanced, with many short turns, while the AI condition looked like supported monologue — fewer, longer turns, a smaller learner share of the floor, and higher within-turn fluency. The AI's verbose, syntactically regular input was associated with greater short-term uptake and stronger syntactic priming after controlling for input volume, and satisfaction tracked learners' own production fluency rather than how much they picked up.

Voice-first design becomes a necessity rather than a convenience when the learner cannot use a screen. Kutti AI (Fadurudeen 2026) inverts edtech's visual assumption to reach an estimated 1.4 million blind children worldwide: children hear curriculum content, answer aloud, and receive spoken feedback with no visual dependency. Three choices carry the design — a lightweight struggle-detection engine fusing response latency, wrong-attempt counts, and keyword hesitation cues to decide when to hint or simplify; a cross-language answer-matching pipeline so that code-switching and pronunciation variation are not penalised; and an offline on-device ASR path that removes the connectivity requirement. It is a systems contribution with no learning-outcomes evaluation, so its pedagogical claims remain hypotheses.

Accent, meanwhile, is a design variable with social consequences. Ravi et al. (2026) had 33 teachers interact with a GenAI voice agent in British, Indian, and African American accent conditions in K-12 group learning. The British-accented agent was treated largely as a tool and engaged with in detached, utility-based ways, while the Indian- and African American-accented agents were more readily anthropomorphized and integrated as peers, with stronger trust and reliance over time. Turn-taking, questioning patterns, and perceived social presence all shifted with accent — the voice that makes an agent easy to treat as a neutral resource can make it easier to ignore as a collaborator.

Heavier use of a spoken support tool is not automatically better. Kuang, Li and Weng (2026) had 22 interpreting trainees complete bidirectional computer-assisted consecutive interpreting tasks in systems built on ASR and machine translation, capturing eye movements, pen notes, and voice output. Four interaction profiles emerged — Intensive Engagers, Fast Scanners, Traditionalists, and Frequent Switchers — and 58.3% of stage-level observations changed profile between comprehending the source and producing the target. Only comprehension-stage patterns predicted product quality, and the AI-heaviest cluster scored lowest on fluency of delivery (5.46 against 6.17–6.46) and target language quality (5.70 against 6.35–6.67). Attention spent reading an AI transcript is attention not spent building one's own representation; the authors recommend teaching learners to reflect on their own strategy rather than prescribing one, since the pattern is invisible unless surfaced.

AI in Oral and Spoken Assessment

Oral assessment has always been pedagogically strong and logistically expensive, and speech technology now attacks the cost side. Pentland, Lowenthal and Krier (2026) describe asynchronous oral assessments (AOAs): prompts delivered just-in-time, brief time-limited webcam responses that cannot be revisited, and grading against embedded rubrics with auto-generated transcripts. Across two courses taught by a single instructor without scheduling constraints, students scored higher on AOAs than on in-person multiple-choice exams — significantly in Study 2 (midterm median 92.5 vs 70, p < .001; final 94.2 vs 86.4, p = .002) and directionally in Study 1, with moderate cross-format correlations (τ = .44; τ = .25, non-significant). Students shifted to more active preparation, 90.91% saw the format as closer to workplace communication than written exams, and 81.82% reported engaging more actively with content. The authors are careful that these are format score differences, not evidence of learning gains, and that the integrity advantage is inferred rather than measured; a re-scoring check with an LLM found instructor scores systematically higher with moderate-to-good agreement (ICC 0.73 and 0.60), offered as a reliability check rather than an endorsement of automated grading.

Fenton (2025) supplies the rationale in review form: because orals are real-time and interactive, students cannot generate answers in advance and memorise them, and the format probes reasoning at higher levels of Bloom's taxonomy rather than recall. Its benefits — personalization, authenticity, work-readiness, deeper knowledge — come with named costs: scheduling and logistics, learner anxiety, and bias risk around gender, ethnicity, language, response speed, and non-anonymous marking, though the evidence suggests orals can be as inclusive as written exams for some learners, including those with dyslexia.

Two deployments show what the technology adds. AkoVoice (Adams 2026) assessed 33 learners across four Level 3 Automotive classes and one Level 3 Engineering class in New Zealand, on learner-owned phones against one mid-range laptop running open-weight models — Mistral 7B via Ollama, faster-whisper for speech-to-text, Chatterbox for TTS — with up to 12 learners assessed simultaneously and no internet connection at any point. AI was cast as evidence-surfacer rather than judge, leaving the assessor's judgment intact. Learner reaction was positive: 21 of 33 (64%) called the task realistic, and none of the 33 disagreed that speaking in real time felt right compared with a written portfolio. The performance data is more interesting. Word counts for identical questions varied five- to eight-fold between learners in every cohort (93–523, 92–796, 309–811), yet verbosity did not predict accuracy on closed questions; nine learners answered in 2 to 13 words and were all marked correct, the shortest being "3500 kgs", while the fuller "the safe working load is three tons" was flagged against a 3.5-ton marking guide. Adjacent Horticulture (14 learners) and Dairy (11 learners) trials found a 95% match between the agent's preliminary grade and the human tutor's grade. On the item side, Aryadoust and Wong (2026) compared iterative prompt engineering against fine-tuning for L2 listening items, producing 40 tests and 240 multiple-choice items: prompt refinement improved quality but plateaued, while fine-tuning GPT-4.1 on the optimized prompt with prompt design held constant improved generation further. Spoken assessment also risks measuring the wrong thing — Morphew et al. (2026) track gesture alongside transcribed speech and argue that assessing only speech misses embodied evidence of understanding, so that fluency can be mistaken for conceptual knowledge.

Accessibility and Equity in the Spoken Channel

Speech technology's accessibility record is mixed: the same channel that removes one barrier can install another. Deaf and hard-of-hearing learners are the clearest mismatch. Chen et al. (2026) designed an LLM question-generation system for DHH learners watching video, drawing on Language Deprivation Theory to explain why text-based prompting fits poorly with sign-based first languages. It added Visual Questions (timestamps where visual information is likely misread — rapid movement, misaligned captions, dense on-screen text) and Emotion Questions (timestamps where prior DHH learners reported frustration or confusion), refining a final bank of 30 questions with learners and instructors. With 16 users the prototype improved Self-Efficacy (M = 5.70, SD = 1.12 on a 7-point scale), and Deaf participants selected visual questions more often than hard-of-hearing participants, who reported reading captions fast enough not to need them. Accessibility here had to be built into generation, not appended to it.

AkoVoice adds two equity considerations that rarely surface together. First, data: recordings were encrypted at rest, hosted locally, and deleted after 90 days, and the authors treat a participant's voice as a personal and cultural expression rather than a resource for training models or building biometric profiles — a stance that matters because the EU AI Act classifies AI used to evaluate learning outcomes in vocational training as high-risk. Second, accent and register: speech-to-text accuracy was not fully tested across learner accents, and rubric criteria such as "engaging" were flagged as culturally specific and potentially penalising to second-language speakers, prompting a review notes field. Set beside the accent findings in K-12 group work and the near-total visual assumption of mainstream edtech, the pattern is that the voice interface removes sight, reading-speed, and connectivity barriers while leaving accent, recognition accuracy, hardware cost, and cultural register open.

Scope check. This page is about the spoken channel and the technology that carries it. For second-language pedagogy at large — writing development, grammatical accuracy, motivation, teacher practice — go to Language Learning, which spans L2 instruction with AI and treats speaking as one strand among several. For text-based chatbots and dialogue design independent of modality, go to Conversational AI; several articles here involve dialogue, but their question is about speech, not about prompting a chat model. For screen readers, captioning, tactile graphics, and assistive tools that are not primarily speech-based, go to Assistive Technology, which sits under the broader Accessibility and Inclusive Learning strands of this wiki.

What Remains Uncertain

Almost every strong claim here rests on one short study. The two TTS lesson studies ran in single Japanese high schools with fixed presentation order that confounded format with content, and their learning-outcome measures were self-reported on in-house scales. The ASR feedback model came from one Chinese university and was cross-sectional, so its proficiency-moderation finding describes covariance rather than causation. The willingness-to-communicate study measured perceptions of feedback, not feedback quality. The interpreting profile study is a 22-participant exploratory analysis in which output-stage effects did not reach significance. The AOA and AkoVoice deployments are single-institution, and AkoVoice reports no summative learning evidence because its voice assessments ran in parallel to actual assessments. Kutti AI is a systems contribution with no outcome evaluation at all. What the collection does support is a set of design claims: feedback quality beats practice volume in ASR use, dialogue narration can aid comprehension at the cost of naturalness, format should be matched to learner characteristics rather than fixed, AI should surface evidence rather than replace the human assessor, and speaking in real time measures something writing does not — including an accuracy surprisingly indifferent to how many words the learner uses. Whether any of that holds at another site, another language, or another year is untested.

Connected Concepts

Connected Articles

Embed this page

Copy the code below to embed a chromeless version of this page in a learning management system or other website. The embedded view hides the site header, navigation, and footer.