Research Article
From tool to scaffold: Structured human–AI collaboration and its effects on academic writing and digital critical thinking
Synthesis: This mixed-methods quasi-experiment tested whether pedagogically structured Human AI Collaboration improves academic writing and digital Critical Thinking among Saudi EFL undergraduates. Over 10-11 weeks, fifty-three students in two intact sections, experimental (n = 31) and control (n = 22), worked a five-part workflow: problem framing and prompt design, iterative drafting, revision, verification of claims and citations, and responsible-use regulation, using a Generative AI assistant and a language-feedback tool; the control group followed the same syllabus without AI integration. At posttest the experimental group scored far higher on both writing and digital critical thinking, and the two outcomes moved together almost in lockstep, with reflections showing AI as a planning and revision scaffold alongside verification routines and ethical self-regulation. Partly self-reported and single-site, the sample makes the effects preliminary upper bounds.
Key Findings
- Fifty-three undergraduates at one Saudi public university, assigned by intact class to an experimental (n = 31) or control (n = 22) group after baseline equivalence was confirmed.
- Posttest writing proficiency was 15.23 (experimental) versus 11.91 (control): a 3.32-point difference (95% CI [2.44, 4.20]), Hedges' g = 2.09, t(51) = 7.60, p < 0.001.
- Posttest digital critical thinking was 89.84 versus 56.18 on the 20–100 scale — a 33.66-point gap (95% CI [28.21, 39.11], Hedges' g = 3.41, p < 0.001) — while the control group barely moved (p = 0.329).
- Writing and digital critical thinking correlated at pretest (r = 0.81), posttest (r = 0.94), and gain scores (r = 0.93), all p < 0.001.
- Four themes emerged from experimental-group reflections: AI as a writing scaffold (100%), detecting and repairing AI errors (100%), verification as routine practice (81%), and metacognitive regulation with ethical self-positioning (100%).
- Students reported verifying AI output before use (M = 4.00, SD = 0.86 on a 1–5 scale) far more than relying on it beyond course guidance (M = 2.10, SD = 0.30 on a 1–3 scale).
How the intervention was designed and measured
The study used a convergent mixed-methods quasi-experimental pretest–posttest control-group design in a required academic writing course. Writing was scored on a 45 min timed essay of 250–300 words using a four-dimension analytic rubric (total 0–20), two raters blind to group. Digital critical thinking used the 20-item Digital Critical Thinking Scale (1–5 items, total 20–100; α = 0.97 pretest, 0.99 posttest).
Over 11 weeks the experimental group received instruction in five components: problem framing and prompt design; generating alternative outlines, arguments, and claims; iterative drafting and revision; verification and triangulation of claims and citations in scholarly databases; and responsible-use regulation covering attribution, reliance, and authorial ownership. Assignments required an AI-use log of prompts, outputs consulted, and accept-or-reject justifications, preserving learner agency.
Writing and digital critical thinking moved together
Within the experimental group, writing rose from 11.19 to 15.23 and DCTS from 57.03 to 89.84 (t(30) = 46.58 and 27.59, both p < 0.001); the control group's writing gain was negligible (0.73 points) and its DCTS did not move (t(21) = 1.00, p = 0.329). All three null hypotheses were rejected. The outcomes grew tighter across the interval — r = 0.81 at pretest, 0.94 at posttest, 0.93 in gain scores (all p < 0.001) — which the authors read as co-development of two competencies rather than one-way causality.
What students described doing with AI
Reflections yielded four themes. Theme 1 cast AI as a flexible writing scaffold: students requested outlines and alternative claims, then rewrote content in their own words anchored in course readings. Theme 2 captured active error repair — fabricated citations and unverifiable statistics, cultural mismatches, and inappropriate register. Theme 3 documented verification as routine critical practice through Google Scholar, library databases, and course materials, keeping only externally supported content. Theme 4 was metacognitive regulation and ethical self-positioning: drafting first without AI, restricting it to grammar and vocabulary, setting time limits, and stressing authorship and avoiding plagiarism. A joint display ties scaffold use to the writing gains and error detection to the critical-thinking gains.
Boundary conditions the authors flag
The authors caution that the effect sizes invite scrutiny: novelty effects, greater instructional intensity in the experimental condition, a self-report instrument (the DCTS) given to students already practiced in what it measures, and item redundancy behind a posttest α of 0.99. Guided AI-supported instruction may backfire under thin Scaffolding, low digital critical thinking or metacognitive readiness, gradual delegation of composing decisions (a reliance risk posttests may not detect), or weak institutional support such as unclear AI-use policies and untrained instructors. The model fits Saudi Vision 2030 and culturally responsive, integrity-aware teaching rather than prohibition.
What this means for practice
- Instructors. Teach the five-step workflow — frame, draft, revise, verify, regulate — and require an AI-use log of what students accepted, rejected, and why; the gains track those routines, not tool access.
- Instructional designers. Have students cross-check AI claims and citations for fabrication against Google Scholar and course readings before citing.
- Faculty developers. Treat digital critical thinking as a taught and assessed target — hallucination detection, source triangulation, authorship boundaries — not a by-product of AI exposure.
- Administrators. Replace blanket bans with explicit authorship, attribution, and disclosure expectations; clear rules preserve authorial control.
- Institutions. Publish AI-use governance making permitted assistance explicit; benefits track institutional clarity, not tool availability.
Limitations
- Assignment by intact class in a single institution, with only two clusters: standard t-tests ignore intraclass clustering and multilevel models were not estimable, so reported p-values may be optimistic.
- N = 53 at one Saudi public university; the authors advise against generalizing beyond comparable Saudi EFL contexts without replication.
- Writing was assessed with one timed task per time point, and digital critical thinking rested entirely on self-report; a posttest α of 0.99 may reflect item redundancy rather than construct breadth.
- The paper does not name the generative AI assistant or the language-feedback tool, or their model versions, and reports no calendar dates for the 11-week intervention; process reflections came only from the experimental group.
Citation
Alshehri, O. A. O., Lamouchi, A., Abdellatif, M. S., Al-Dosari, M. N. A., Mekheimer, M., & Nemt-allah, M. A. (2026). From tool to scaffold: Structured human–AI collaboration and its effects on academic writing and digital critical thinking among Saudi EFL learners. Frontiers in Psychology.