Research Article
A Semi-Automated System for Generating Dialogue-Based TTS Lessons Using Large Language Models: An Exploratory Study of Educational Potential
Synthesis: This study presents a three-stage, human-in-the-loop pipeline in which LLMs generate slides and TTS-optimized narration (including an Expert-×-Novice dialogue format inspired by cognitive apprenticeship) while educators retain review authority at each stage. In a classroom quasi-experiment with 245 first-year high school students, replacing instructor voice with TTS audio did not degrade the core learning experience, and the dialogue format was rated significantly higher than single-speaker TTS on comprehension, cognitive engagement, and preference — at the cost of less natural audio.
Overview
High-quality lesson video is a prerequisite for models such as the flipped classroom, yet producing it demands substantial time and specialized skill, which raises the barrier to entry for many educators. Generative AI and text-to-speech LLMs and TTS technology promise to lower that barrier, but existing automated content-generation systems remain at the level of technical evaluation and have not been validated against the educational experience of actual learners. Moreover, dialogue-based learning research has historically presupposed a human teacher.
This work addresses the gap by combining three concerns that had been treated separately: LLM-based automation of lesson production, the educational acceptability of TTS audio compared with a human voice, and the design of dialogue-format narration. It is positioned as an empirical study of Generative AI in real classroom settings, deliberately designed to augment rather than replace educators.
Study Design & Method
System (three-stage, human-in-the-loop). The proposed pipeline keeps the educator in final control at every stage, explicitly rejecting a "one-shot" fully automated tool. Stage 1 uses Claude 3.5 Sonnet with a slide-design optimization prompt to turn unstructured input (textbook chapters, lecture materials) into Marp Markdown slides; the educator fact-checks, refines the pedagogical flow, and adds figures and tables (Human Feedback 1). Stage 2 generates narration scripts optimized for TTS "listenability," with nine optimization strategies and two selectable modes: single-speaker (structured, one-way delivery with linguistic emphasis cues) and dialogue. The educator verifies accuracy, clarity, and alignment with teaching style (Human Feedback 2). Stage 3 synthesizes speech via the Gemini TTS API (gemini-2.5-pro-preview-tts), renders Marp slides through HTML and Selenium, and integrates audio and images with MoviePy; audio is 24 kHz mono 16-bit PCM WAV and inline prosody tags control pauses.
Dialogue design. The novel contribution is automatically generating Expert-×-Novice dialogue narration: an Expert persona supplies accurate knowledge through metaphors and concrete examples while a Novice gives voice to learners' thought processes, following a five-stage dialogue pattern. The design is inspired by cognitive apprenticeship theory but does not faithfully implement all of its elements — notably, fading is not realized, so the authors deliberately use the phrase "inspired by" throughout.
Evaluation. To validate the system's utility, the authors ran a quasi-experiment with 245 first-year students at a prefectural high school in Niigata Prefecture, delivered as part of an inquiry-based, data-driven career education curriculum using generative AI. The same cohort sequentially experienced three formats: (1) 21 May 2025, traditional video with instructor voice (control condition); (2) 25 June 2025, single-speaker TTS lesson generated by the system (Experiment 1); and (3) 23 July 2025, dialogue TTS lesson (Experiment 2). This within-subject design had a critical constraint: the order of implementation was fixed and lesson content differed across sessions (career paths; how generative AI works; a framework for data-driven decision-making), so format effects cannot be cleanly separated from order and content effects.
Post-lesson questionnaires on 5-point scales yielded 229 valid responses (93.5%) for the single-TTS lesson and 206 (84.1%) for the dialogue lesson. Because responses were collected anonymously, the two rounds cannot be individually matched. The instrumentation operationalized three frameworks — the ARCS model (Attention, Relevance, Confidence, Satisfaction), Cognitive Load Theory (extraneous and germane load), and learning engagement theory (emotional and cognitive) — with all items created in-house for this study. Two analyses followed: a within-subject retrospective comparison of the three formats included in the Experiment 2 questionnaire (N ≤ 183, Friedman test plus TOST equivalence testing with bounds Δ = ±0.5), and a repeated cross-sectional between-group comparison of detailed items (Mann–Whitney U with BH-FDR correction).
Key Findings
- Prior-knowledge imbalance favored the comparison group, not dialogue. 66.0% of the dialogue TTS group reported knowing "nothing at all" about the content vs. only 34.1% of the single TTS group (χ²(1) = 43.05, p < .001), while subjective difficulty did not differ (U = 24,959, p = .270). The authors treat this as a confound that systematically disadvantages the dialogue group, making their advantage a conservative estimate.
- RQ1 — TTS did not substantially degrade the learning experience. The Friedman test found no significant difference across the three formats on comprehension, concentration, or overall evaluation (p > .14 for all three metrics, effect sizes r < 0.11). TOST equivalence testing confirmed all three core metrics and all pairwise comparisons fell within the ±0.5 equivalence range (p < .0001).
- RQ2 — dialogue TTS was significantly superior on comprehension and cognitive engagement. Comprehension: p = .006, q = .025, r = .130. Cognitive engagement ("deepened thinking"): p = .019, q = .048, r = .121. "Can explain to a friend": p = .004, q = .021, r = .152.
- Enjoyment missed correction but survived a covariate model. Emotional engagement ("enjoyable") did not reach significance after FDR correction (q = .081), yet a proportional-odds model with Prior Knowledge as a covariate confirmed a significant dialogue advantage (OR = 1.65, q = .025), with comprehension, explainability, and cognitive engagement advantages shown not to be attributable to prior-knowledge imbalance.
- There is a trade-off: dialogue audio is less natural. Single TTS was rated significantly more natural in audio (p < .001, q < .001, r = −.238), consistent with increased extraneous cognitive load from multiple speakers — while germane load showed no significant difference between formats.
- Learners preferred the dialogue format. Dialogue received 66.9% support as "most enjoyable to learn from" (χ²(2) = 78.09, p < .001); on perceived highest learning effect, dialogue led at 47.3% vs. 34.0% for instructor video and 18.7% for single TTS (χ²(2) = 18.52, p < .001); 50.5% would like to experience it again (multiple-choice, reference only).
What this means for practice
- Instructors. Replace instructor-voice video with TTS narration where production time is the bottleneck. In a quasi-experiment with 245 first-year high school students, the three formats were statistically equivalent on comprehension, concentration, and overall evaluation (Friedman p > .14 for all three metrics; TOST equivalence within Δ = ±0.5, p < .0001).
- Instructors. Choose dialogue-format narration when comprehension is the priority: dialogue TTS beat single-speaker TTS on comprehension (p = .006, q = .025, OR = 2.24), cognitive engagement (p = .019, q = .048), and being able to explain the content to a friend (p = .004, q = .021). Select the mode to match the objective — keep single-speaker narration where clarity and low cognitive load dominate — and do not expect the format to implement cognitive apprenticeship on its own, since fading is not realized in the generated dialogue.
- Designers. Accept the trade-off knowingly: single TTS was rated significantly more natural (p < .001, r = −.238), so reserve the Expert-×-Novice dialogue format for material where its framing does work that audio polish cannot, and offset the multiple-speaker load with the pauses, subtitles, and speaker-name overlays the authors recommend.
- Software developers. Keep the educator review gate at every one of the three pipeline stages — slide generation, narration scripting, and TTS rendering — because the system's value rests on staged human supervision rather than one-shot automation, a Human-in-the-Loop pattern for instructional video rather than an argument for removing teachers. Treat listenability as an engineering requirement: optimize the script at the prompt level, because human-written narration uses written-language structures that TTS renders awkwardly, and measure the instructor effort each stage costs.
- Software developers. Check prior-knowledge balance before claiming a format effect: 66.0% of the dialogue group reported knowing "nothing at all" about the content versus 34.1% of the single-TTS group (χ²(1) = 43.05, p < .001), an imbalance the authors treat as making the dialogue advantage conservative. Pilot the pipeline with further populations, such as university students and vocational trainees, before generalizing beyond first-year high schoolers.
Limitations
- Format effects cannot be separated from content and order effects: the same cohort experienced three lessons in a fixed order with different content (May 21, June 25, and July 23, 2025), and the authors attribute the "useful in the future" difference to content rather than format.
- Questionnaires were collected anonymously, so the single-TTS (N = 229) and dialogue (N = 206) rounds cannot be matched to individuals; independence of observations is not guaranteed, selection bias cannot be ruled out, and the repeated cross-sectional comparison is treated only as an approximation.
- Measurement covered immediate subjective evaluation alone: knowledge retention, transfer, and sustained viewing behavior were not measured, and no pre-/post-tests were run.
- All items were created in-house with reliability coefficients, factor analysis, and scale validity left unverified, so results are self-assessed differences on individual items rather than construct-level changes such as ARCS attention or germane load.
Citation
Kumoi, G., Watanabe, F., Suko, T., Ishida, T., et al. (2026). A Semi-Automated System for Generating Dialogue-Based TTS Lessons Using Large Language Models: An Exploratory Study of Educational Potential.