On this page

Synthesis: Chen, Zeng, Ou, and Zhong (2026) surveyed Chinese undergraduates practicing English speaking with Automatic Speech Recognition (ASR) tools and modeled how the technology's features shape learning. They find that the quality of ASR-based Feedback and well-designed reflection tasks drive both feedback internalization and reflective behavior, while frequent use chiefly boosts reflection rather than motivation, and recognition accuracy raises intrinsic Motivation without directly triggering deeper reflection. Both reflective behavior and motivation predict speaking gains, but language proficiency strongly moderates these effects: more proficient learners internalize and benefit from ASR feedback far more effectively, while weaker learners need extra support.

Key Findings

  • Feedback quality, not quantity, drives learning. Accurate error correction and structured reflection tasks significantly improve feedback internalization and reflective behavior. Simply using the system more often does not substitute for pedagogically sound feedback.
  • Usage frequency boosts reflection, not motivation. Frequent ASR use moderately stimulates reflective behavior but has no independent effect on learning motivation — the assumption that mere exposure raises motivation is not supported.
  • Recognition accuracy affects motivation but not reflection. Accurate error detection enhances intrinsic motivation, likely by building confidence and trust in the tool, yet it does not on its own trigger deeper cognitive engagement. Learners may note a flagged error without analyzing its cause.
  • Reflective behavior and motivation both predict speaking gains, with reflection the stronger driver. The cognitive work of self-review translates into oral improvement more directly than affective motivation alone, though both are needed.
  • Language proficiency is a pivotal moderator. For both reflective behavior and motivation, the payoff is weak and non-significant at low proficiency and grows markedly stronger at average and high levels. Less proficient learners likely lack the foundational linguistic knowledge to act on ASR feedback and may need simplification, explicit strategy guidance, and extra scaffolding.

Study Design & Method

The study surveyed undergraduates at a teacher-training university in China who practice English speaking with ASR tools. Of 400 questionnaires distributed (half online, half offline), 325 valid responses were retained. Data were analyzed with exploratory and confirmatory factor analysis plus structural equation modeling to test how ASR accuracy, usage frequency, corrective feedback quality, and reflection task design influence reflective behavior and intrinsic motivation, and in turn oral proficiency improvement. Language proficiency was then tested as a moderator using multi-group structural equation modeling across beginner, intermediate, and advanced learners.

Implications for AI in Education

For Language Learning and Intelligent Tutoring, the findings show ASR's value rests on pedagogical integration, not the tool alone. Educators should prioritize feedback quality over practice volume, scaffold reflection through structured tasks such as recording-and-comparing and reflective journals, and differentiate support by proficiency level — since one-size-fits-all ASR practice risks widening achievement gaps. Developers should build systems that explain errors articulatorially (e.g. where to place the tongue) rather than merely flagging mistakes, and add adaptive features that simplify feedback and increase complexity as learners progress — connecting directly to Self-Regulated Learning and Personalized Learning design. Because accurate feedback can momentarily discourage lower-proficiency learners, systems should pair precision with supportive, positively framed scaffolding.

Limitations

  • Cross-sectional design precludes causal inference; experimental and longitudinal designs are needed to confirm the pathways.
  • Single-university Chinese sample with predominantly beginner-level learners limits generalizability to other cultural and educational contexts.
  • ASR treated partly as monolithic — the unique contributions of specific design elements are modeled but not isolated experimentally.
  • Self-reported measures of proficiency, motivation, and speaking gains, with proficiency grouped into coarse CEFR-based bands, may introduce measurement imprecision.

Connected Concepts

Connected Articles

Citation

Chen, D., Zeng, L., Ou, W., & Zhong, S. (2026). ASR Technology in College English Speaking Instruction: The Role of Feedback Internalization and Metacognitive Strategies. Frontiers in Psychology, 17, 1847238.

Embed this page

Copy the code below to embed a chromeless version of this page in a learning management system or other website. The embedded view hides the site header, navigation, and footer.