On this page

Synthesis: Lu and colleagues (2026) ran a nine-week GenAI-supported opinion-writing program with 301 Grade 5 and 6 students in Eastern China, randomly assigning eight intact classes to the program or to conventional instruction. The program raised students' ideal L2 writing self, academic buoyancy, and behavioural and emotional engagement, and improved language use — but it did not move growth mindset, cognitive or metacognitive engagement, or organization, and the effects were small. The value of the study lies in that pattern: it shows which dimensions of young learners' writing development a GenAI program can shift, and it names the risks — over-reliance, shortcut-seeking, and diminished self-monitoring — that come with it.

Overview

Most research on generative AI in second-language writing has been conducted with university students. Upper-primary learners are a distinct case: they are still building the writing strategies and self-regulation that older students are assumed to have, so a tool that produces fluent text can scaffold their development or bypass it, depending on how it is used. The paper's rationale is that this population needs developmentally informed investigation rather than extrapolation from higher education.

The study therefore asks what changes when a GenAI-supported writing program is embedded in a primary English curriculum, and how students experience those changes. It examines three groups of constructs: writing Motivation (growth mindset, ideal L2 writing self, academic buoyancy), engagement (cognitive, behavioural, emotional, metacognitive), and writing performance across opinion and content, organization, and language use. Using mixed methods is a deliberate choice: the quantitative design establishes which constructs moved, and interviews explain the processes behind the pattern.

Study Design & Method

  • Participants and design. Four intact classes per grade in Grades 5 and 6 at one primary school in Eastern China. Classes were randomly assigned at the class level to GenAI-supported or conventional instruction to avoid contamination between conditions: the experimental group had 151 students (75 Grade 5, 76 Grade 6) and the control group 150 (74 Grade 5, 76 Grade 6). The final analytic sample was 301 students — 149 in Grade 5 and 152 in Grade 6, 47.5 percent male, ages 10 to 13 (M = 10.82, SD = 0.75). Consent came from students, parents or guardians, and instructors.
  • Baseline equivalence. Because assignment was by class, the authors checked baseline writing performance using the most recent grade-level standardised English exam. The two groups were statistically indistinguishable (experimental M = 5.23, SD = 0.92; control M = 5.21, SD = 0.99; t(299) = 0.175, p = .86, d = 0.02), and gender and age were balanced (all p > .20). Grade-level baselines were also compared after standardising across the full sample (Grade 5 M = −0.09; Grade 6 M = 0.09; t(299) = 1.55, p = .12, d = 0.18), which supported pooling the grades. Both grades wrote opinion prompts scored with the same rubric.
  • The program. Nine weeks, one 40-minute session per week, integrated into the school-based curriculum, delivered with standardised lesson plans under both conditions. Week 1 served as kickoff and pre-test: students received a printed handbook with exemplar opinion texts, an analytic rubric and a GenAI user guide containing operational instructions, ethical-use principles and a categorised bank of sample prompts for idea development, structure analysis, language enhancement, quality evaluation, feedback and revision. Weeks 2–3 focused on expressing and supporting opinions, Weeks 4–5 on text structure through GenAI-supported analysis, paragraph reordering and semi-guided writing with AI feedback, Week 6 on collaborative draft revision for language, Week 7 on comparing teacher and AI feedback on anonymised drafts, and Week 8 on revisiting pre-test drafts for peer and AI evaluation followed by individual revision. Students worked in small groups with shared tablets in the collaborative weeks and individually with one tablet each during feedback-intensive weeks. Prompting was bilingual (Chinese and English) and pitched at Grade 5/6 language level.
  • Measures and analysis. Motivation, engagement and performance were assessed at pre-test (Week 1) and post-test (Week 9). One-way MANCOVA tested between-group differences controlling for pre-test scores, with partial η² for effect size, followed by paired-samples t-tests for within-group change (Cohen's d, Bonferroni-adjusted α = .017). Semi-structured interviews were conducted with 12 students from the experimental group, sampled across high, medium and low pre-test writing groups (interviewed subsample pre-test M = 5.20, SD = 0.96), and thematically analysed with counts of how many interviewed students reported each observation.

Key Findings

  • Motivation improved selectively. The multivariate group effect was significant, F(3, 294) = 3.067, p = .028, Wilks' Λ = 0.970, partial η² = 0.030. Follow-up analyses showed the experimental group outperformed controls on ideal L2 writing self (F(1, 296) = 6.707, p = .010, partial η² = 0.022, adjusted Mdiff = 0.20 [95% CI 0.05, 0.35]) and academic buoyancy (F(1, 296) = 5.593, p = .019, partial η² = 0.019, adjusted Mdiff = 0.17 [95% CI 0.03, 0.32]). Growth mindset showed no significant difference (p = .239).
  • Within-group gains followed the same pattern. In the experimental group, ideal L2 writing self rose (Mdiff = 0.24, t(150) = 3.218, p = .002, d = 0.26) and academic buoyancy rose (Mdiff = 0.20, t(150) = 2.90, p = .004, d = 0.24), both surviving Bonferroni correction. Growth mindset did not (p = .079, d = 0.14), and the control group's modest growth-mindset gain (p = .018, d = 0.20) did not survive correction.
  • Engagement moved in two of four dimensions. The engagement MANCOVA was significant, F(4, 292) = 8.392, p < .001, Wilks' Λ = 0.897, partial η² = 0.103, with group effects on behavioural engagement (F(1, 295) = 3.910, p = .049, partial η² = 0.013) and emotional engagement (F(1, 295) = 4.310, p = .039, partial η² = 0.014). Cognitive engagement did not reach significance (p = .056, partial η² = 0.012), nor did metacognitive engagement (p = .492). Within-group gains appeared for emotional (Mdiff = 0.23, t(150) = 2.721, p = .007, d = 0.22) and behavioural engagement, but not cognitive (p = .157, d = 0.12) or metacognitive (p = .390, d = 0.07).
  • Performance improved on language use, less clearly on opinion and content. The performance MANCOVA was significant, F(3, 294) = 2.841, p = .038, Wilks' Λ = 0.972, partial η² = 0.028, with a group effect on language use (F(1, 296) = 4.246, p = .040, partial η² = 0.014) and a near-significant effect on opinion and content (p = .054, partial η² = 0.013) tending to favour the experimental group. Organization (p = .898) and total scores (p = .104) showed no group differences.
  • Students' accounts explain the motivational shift. Eight of the 12 interviewees said that seeing GenAI produce fluent text with vocabulary they already knew made successful writing feel attainable — "many of the words were ones we had already learned… seeing its writing made me feel that I could also write a good piece" (S11, low group). Six reported a sense of achievement from concrete improvements, and five described a stronger awareness of authorship: S1 (high group) reported rewriting the handbook prompts in their own words, selecting and revising ideas rather than copying output, a claim the authors triangulated against the recorded interaction and draft comparisons.
  • Academic buoyancy showed up as manageability. Nine interviewees described the process as smoother and described more confidence with tasks they had previously found difficult.
  • The authors report cautionary roles alongside the gains. Their synthesis lists five risks: reinforcing reliance on the tool instead of personal effort, shortcut-oriented strategies that hinder gradual skill development, constrained independent Problem Solving, reduced use of previously taught writing strategies, and diminished self-monitoring during writing.

Implications for AI in Education

The paper's most usable lesson is that GenAI support does not lift a construct like "writing development" uniformly — it moves specific dimensions and leaves others untouched, and the pattern is diagnostic. Motivation improved through the aspirational self and resilience rather than through growth mindset, which suggests that simply handing students an impressive tool does not teach them that ability grows with effort; the authors recommend designing activities around the sequence of challenge, effort and progress so that the tool reads as support rather than shortcut. Likewise, engagement improved emotionally and behaviourally while cognitive and metacognitive engagement did not, which is a warning that enjoyment and activity are not the same as deeper processing, and that students may be doing less self-monitoring when the tool is available.

For Scaffolding and assessment design, the program's structure is instructive. It paired a Feedback-rich sequence — comparing teacher and AI feedback on anonymised drafts, revisiting pre-test drafts for peer and AI evaluation, individual revision — with explicit instruction in how to prompt, including a bank of prompts tied to writing goals. Where writing education aims at discourse-level skills such as organization, this study offers little comfort: those were exactly the dimensions that did not move, so tasks that use AI to analyse and reorder structure still need deliberate design if organisation is the target.

The authors also flag how much the setting mattered. Students had routine access to GenAI-enabled tablets and were already fluent with digital tools, and teachers provided real-time support, so participation required no additional training. Where infrastructure, familiarity or teachers' pedagogical awareness of AI-supported instruction are lower — particularly in many developing countries — the authors expect substantially different results, and they ask that findings be read as context-specific rather than globally generalisable.

Limitations. The sample came from upper grades in a single primary school in China. Motivation and engagement were measured mainly by self-report, which may not capture observable classroom behaviour; the authors suggest classroom observation or multi-method designs. Both Grades 5 and 6 were included but developmental differences between the grades were not examined, and the study analysed outcome change without modelling the relationships among motivation, engagement and performance over time.

Connected Concepts

Connected Articles

Citation

Lu, Q., Wang, M., Yao, Y., Zhu, X., Xiao, L., & Yin, H. (2026). Exploring the impact of a GenAI-supported writing program on primary students' writing motivation, engagement, and performance: A mixed methods study. Learning and Instruction, 106, 102451.

Embed this page

Copy the code below to embed a chromeless version of this page in a learning management system or other website. The embedded view hides the site header, navigation, and footer.