On this page

Synthesis: Lu et al. (2027) use an eight-week intervention with 60 Grade 5 students in a Chinese primary school to examine how conversational AI bots support critical thinking during multimodal writing. Students transformed their narratives into AI-generated images and short videos, and the study tracked six Critical Thinking dimensions at pre-, post-, and one-month-follow-up. Gains were uneven: interpretation, analysis, evaluation, and explanation rose and held, self-regulation improved only short-term, and inference did not change — suggesting the multimodal externalization of meaning can reduce the inferential demand that writing usually carries.

Key Findings

  • Dimension-specific rather than uniform gains. Repeated-measures ANOVA and Bonferroni-adjusted comparisons showed sustained increases (T1→T2 and T1→T3) in interpretation, analysis, evaluation, and explanation, with no significant change between T2 and T3. Self-AI Regulation in Education rose significantly from T1 to T2 but not to the one-month follow-up (T1→T3 n.s.), and inference showed no significant change across any comparison.
  • Multimodal externalization as a scaffold for reflection. AI-generated images of students' own narratives functioned as visual cues that prompted three interlinked processes — idea verification, idea gap detection, and idea revision. Seeing their stories rendered visually helped students interpret meaning, notice unclear or incomplete parts, view their writing from a reader's perspective, and evaluate whether AI-generated details were reasonable (e.g., catching an "impossible" bike-riding image).
  • Resemiotisation and cross-modal comparison. Framed through resemiotisation theory (meaning reconfigured across modes), the image-and-video composing let students compare written, visual, and auditory representations, supporting interpretation, analysis, and explanation. Combining images with recorded narration added a "director's perspective" that surfaced logical and sequencing problems.
  • Explicit visuals may reduce inference opportunities. A key cautionary finding: because images made story details explicit, students reported less need to infer implicit meanings from text alone — "with the pictures, it felt easier to understand the story without guessing too much." The authors treat this as a possible explanation (not proof, given the small sample) for inference's flat trajectory.
  • Peer collaboration supplied complementary inference. During revision, peer questions prompted students to infer why classmates misread parts of the story and how other readers might interpret the same scenes — an occasion for inference the solo AI interaction did not provide, alongside opportunities to evaluate peer suggestions and reflect on revision decisions.

Synthesis

The study contributes a rare upper-primary, multimodal test of conversational AI bots as learning partners in writing. Its core insight is a trade-off: externalizing meaning across visual and audiovisual modes scaffolds several Critical Thinking dimensions and suits young learners whose cognitive and metacognitive capacities are still developing, but it can simultaneously lower the inferential load that text-only writing imposes. The design implication is that multimodal AI composing should be paired with continued Scaffolding and structured peer collaboration that deliberately restore occasions for inference and sustain self-regulatory reflection, rather than relying on the bot alone. Methodologically, dimension-level analysis — not an aggregate critical-thinking score — was essential to revealing this uneven pattern.

Connected Concepts

Connected Articles

Citation

Lu, Q., Yao, Y., Zhu, X., Xiao, L., & Yin, H. (2027). More externalization, but less inference? Exploring changes in young learners' critical thinking during conversational AI bot-supported multimodal writing practice. Computers in Human Behavior, 186, 109141.

Connected FAQs

Embed this page

Copy the code below to embed a chromeless version of this page in a learning management system or other website. The embedded view hides the site header, navigation, and footer.