Yueqiao Jin, Kaixun Yang, Roberto Martinez-Maldonado, Dragan Gašević & Lixiang Yan (2026) — Computers and Education: Artificial Intelligence (Elsevier), Article in Press. Open Access, CC BY 4.0. doi:10.1016/j.caeai.2026.100655.
📄 Full text (ScienceDirect, OA)
Summary
A randomized experiment (n = 79 medical/nursing students) examining how the initiative design of an AI writing agent shapes reasoning, agency, and immediate independent performance. Students completed two multimodal analytical writing tasks (interpreting healthcare-simulation data visualisations: bar chart, network diagram, ward heatmap) with either a reactive agent (responds only when prompted, n = 39) or a proactive agent (initiates sequenced questions and feedback, n = 40). GenAI literacy was measured with the validated 20-item GLAT. The study introduces the agency gap: a relational mismatch between the initiative an AI agent demands and the learner's capacity to initiate, monitor, evaluate, and internalise AI-supported reasoning — neither an individual deficit nor a fixed property of the system.^[raw/papers/caeai-2026-agency-gap-ai-writing.md]
Key findings
RQ1 — Epistemic network structure differs strongly by design
- ENA (explaining 38.3%/22.7% and 27.5%/26.4% of variance) separated conditions with large effects (Cliff's δ = −0.56, −0.78; both p < .001).
- Proactive dialogues: stronger links between conceptual reasoning, adequate reasoning, and constructive engagement (EP-CS–EP-CP-Adeq, EP-PS–EP-CP-Adeq, I-CON–EP-CP-Adeq) — more integrated epistemic elaboration.
- Reactive dialogues: more factual/procedural/off-task pairings (EP-PS–EP-OFF, EP-OFF–I-ACT) — the learner's own regulation is more visible but discourse stays descriptive.
- The difference is in how ideas are connected, not how often categories appear and not the final score.^[raw/papers/caeai-2026-agency-gap-ai-writing.md]
RQ2 — GenAI literacy predicts immediate independent performance
- After AI support was removed, GLAT predicted Visual Data Integration (OR 1.14, p = .039), Critical Thinking (OR 1.15, p = .029), and the Composite score (OR 1.11, p = .032) — modest, higher-order effects; not significant for insightfulness, organisation, or linguistic quality.
- AI-supported performance strongly predicted AI-removal performance on all dimensions (all p ≤ .001) — continuity, but cannot distinguish learning from stable competence.
- No significant condition effect and no significant literacy-by-design interaction.^[raw/papers/caeai-2026-agency-gap-ai-writing.md]
RQ3 — Mediation patterns are suggestive, not confirmatory
- Reactive condition: significant total literacy→performance association (β = 0.172, p = .020) with direct path remaining (β = 0.140); indirect effect non-significant (95% CI [−0.022, 0.108]).
- Proactive condition: total and direct coefficients near zero; indirect non-significant.
- Pattern is consistent with smaller literacy-related performance differences under proactive scaffolding, but does not establish compensation, mediation, or moderation — hypothesis-generating for future adequately powered tests.^[raw/papers/caeai-2026-agency-gap-ai-writing.md]
RQ4 — Three design heuristics from learner reflections
1. Sustain autonomy through contextual and confirmatory feedback (reactive strength: confirms interpretations, lowers barrier, but redundant for proficient learners). 2. Promote integrative reasoning and immediate independent application through dialogic scaffolding (proactive strength: connects evidence across visuals, prompts self-correction; risks over-scaffolding easy tasks). 3. Ensure equity through adaptive alignment of initiative with learner expertise and task complexity — a uniform interaction style may under-support some learners while over-directing others.^[raw/papers/caeai-2026-agency-gap-ai-writing.md]Interpretation
- Process ≠ outcome: agent design produced large differences in the relational organisation of dialogue but no significant direct effect on immediate writing scores — the mechanism is how epistemic work is distributed, not output quality.
- The agency gap frames the failure modes: under-support (low literacy × strongly reactive design) and over-direction (high capability × rigidly proactive design), echoing scaffolding's expertise-reversal effect and adaptive-scaffolding accounts.
- Practice: make initiative visible and adjustable (request/skip/pause prompting), structure proactive prompts to orient–interpret–connect–synthesise rather than supply answers, and fade prompts as learners demonstrate independence; teach GenAI literacy as part of academic writing (ai-literacy, agentic-ai).
- Limitations: n = 79 underpowered for mediation; medical/nursing sample; immediate AI-removal task measures near transfer, not durable learning; agency gap theorised, not directly measured; no manipulation-check coding of agent turns.^[raw/papers/caeai-2026-agency-gap-ai-writing.md]
Related Pages
- feedback-futures-genai — Agency vs dependency as a core tension in GenAI feedback
- chatgpt-feedback-engagement-genai — Student-side agency, prompting, and metacognitive engagement
- learner-centered-feedback-ai — Teacher-side calibration of AI initiative/voice
- ai-literacy — GLAT-measured literacy as the learner-side capability
- agentic-ai — Agent initiative as a design variable
- scaffolding — Proactive designs as scaffolding; expertise-reversal risk
- over-reliance — Over-direction and dependency risks of proactive designs
- student-experience — Learner perceptions of reactive vs proactive support
- writing-education — Multimodal analytical writing as the task context
- higher-ed — Deployment context (medical/nursing education)
Citation
APA: Jin, Y., Yang, K., Martinez-Maldonado, R., Gašević, D., & Yan, L. (2026). The agency gap in AI-supported writing: How reactive and proactive agent designs shape multimodal reasoning. Computers and Education: Artificial Intelligence. Advance online publication. https://doi.org/10.1016/j.caeai.2026.100655