On this page

Synthesis: Twenty-two Chinese-native interpreting trainees completed bidirectional computer-assisted consecutive interpreting (CACI) tasks in AI-enabled systems built on automatic speech recognition and machine translation, while eye-tracking, pen-recording and voice-recording captured how they actually divided attention between the AI output and their own note-taking. Cluster analysis of 264 stage-level observations produced four interaction profiles — Intensive Engagers, Fast Scanners, Traditionalists and Frequent Switchers — and 58.3% of observations shifted profile between the comprehension and production stages of the same task. Comprehension-stage patterns, but not production-stage patterns, were significantly associated with interpreting quality, and the AI-heaviest profile consistently scored lowest. The study shows that student-AI interaction in a demanding bilingual task is heterogeneous, unstable across task phases, and not automatically beneficial.

Key Findings

  1. Four interpretable interaction profiles emerged from the eye-tracking, note-taking and switching measures: Intensive Engagers (heavy AI reliance, AI reading proportion 0.83), Fast Scanners (scanning-based processing, reading proportion 0.27, highest fixation frequency at 218.6), Traditionalists (minimal AI reliance, reading proportion 0.12, deep fixation proportion 0.56) and Frequent Switchers (shifts between AI support and handwritten notes, switch count 36.0 against 3.3 to 7.7 in the other clusters).
  2. Interaction patterns were not stable traits: the overall stage-to-stage transition rate was 58.3%, so most students reconfigured how they used the AI between comprehending the source speech and producing the target speech.
  3. Prior AI training mattered: targeted CACI training was associated with more stable, AI-oriented patterns, while untrained participants showed greater reconfiguration across stages.
  4. Only comprehension-stage patterns predicted product quality. Intensive Engagers scored significantly lower on fluency of delivery (5.46 vs 6.17, 6.08, 6.46) and target language quality (5.70 vs 6.35, 6.28, 6.67) than all three other clusters; information completeness and temporal speech measures did not differ by cluster.
  5. The authors caution that the four profiles are context-bound behavioral patterns nested within only 22 participants, not stable learner types, and that the pattern differences for output-stage data were not statistically significant.

Why the heaviest AI users did not perform best

The counter-intuitive finding is that the cluster reading AI output most (AI reading proportion 0.83 at the comprehension stage) produced the weakest delivery and language quality. The authors interpret this as critical engagement mattering more than access: attention spent on the AI subtitle stream is attention not spent on building a personal representation, and the Traditionalists and Fast Scanners retained more of their own processing route. Their recommendation is not to withhold the tool but to teach students to describe and reflect on their own interaction strategy, because the pattern is implicit — hidden in where the eyes go — and therefore invisible to the learner unless it is surfaced.

The study also makes a metacognitive argument for training: since patterns shift as task demands change, students need adaptive expertise — the ability to judge, at each stage, whether the AI support or their own note-taking is the better resource. Instructors are advised to avoid prescribing one correct way of working with AI and instead to guide students toward a strategy they can justify and revise.

Method and measures

Participants performed bidirectional tasks in a CACI environment that integrated ASR and MT features. Behavioral data were triangulated from three channels: eye-tracking (AI reading proportion, regression rate, saccade amplitude, fixation frequency, deep fixation proportion, switch count), pen-recording (note-taking duration and note counts, computed separately for the AI-support area and the notepad area) and voice-recording (target speech duration, pause count, pause duration proportion, fillers and disfluencies). Interpreting quality was rated by two experienced interpreters on an 8-point scale across four bands for information completeness, fluency of delivery and target language quality; measures were computed separately for the input and output stages of each task.

Implications for AI-assisted language and interpreting instruction

  • Treat the AI readout as one resource among several rather than the default: students who split attention across AI output and their own notes performed better on delivery and target-language quality.
  • Make strategy visible: reflection and self-report exercises are needed precisely because interaction patterns are not introspectable.
  • Train for stage-specific strategy, since the same student may benefit from AI support while comprehending and from independent production while speaking.
  • Be cautious about generalizing from a 22-participant exploratory study, especially where output-stage effects did not reach significance.

Connected Concepts

Connected Articles

Citation

Kuang, H., Li, J., & Weng, Y. (2026). Student-AI interaction in computer-assisted consecutive interpreting: patterns and performance. Frontiers in Psychology, 17, 1890808.

Embed this page

Copy the code below to embed a chromeless version of this page in a learning management system or other website. The embedded view hides the site header, navigation, and footer.