Research Article
Rethinking AI-assisted writing instruction: feedback literacy scripts, calibration training, and student writing development
Synthesis: In AI-assisted writing, the value of generative feedback depends less on its abundance than on whether students can evaluate it, calibrate their own self-judgment, and turn external support into independent revision. Dai (2026) shows, in a 2 × 2 factorial experiment with 120 undergraduate English majors, that feedback-literacy training (a Feedback Literacy Script, FRAC) primarily improves writing quality, effective feedback uptake, and deep revision, whereas calibration training (an Assessment-Performance Calibration Activity, APCA) primarily improves Self-Assessment accuracy and reduces overconfidence. Combining the two produced the highest writing gains and strongest retention after AI support was withdrawn, but did not outperform APCA alone on self-assessment accuracy.
Core Finding
Background and Motivation
Generative AI tools such as ChatGPT and Claude now deliver immediate, repeated, and abundant writing feedback across grammar, vocabulary, structure, and argumentation. This changes the learner's task: rather than merely accessing feedback, students must judge whether AI suggestions are accurate, relevant, and worth acting on. The author argues the central pedagogical challenge is no longer access but the capacity to evaluate, select, and implement feedback critically. Prior literature treated AI mainly as an input condition (AI vs. non-AI) and largely overlooked the internal processes through which students judge feedback credibility, reassess their own texts, and translate judgments into revision — a gap this study addresses with process data and a framework linking feedback use, self-assessment calibration, revision behavior, and writing improvement.
The Two Interventions
FRAC is a Feedback Literacy Script targeting the processing of external feedback. Its four steps are Filter (identify the problem type addressed), Reason (judge validity and relevance to the task), Act (implement the feedback as a specific revision), and Check (review whether the revision achieved the intended result). APCA is an Assessment-Performance Calibration Activity targeting internal monitoring, organized around "self-assessment–comparison–adjustment" with four stages: Anchor (review sample texts at different quality levels), Predict (estimate one's own score and confidence), Compare (detect bias against teacher ratings), and Adjust (revise evaluation strategies across rounds). FRAC is grounded in feedback literacy research (Carless & Boud), APCA in calibration and self-regulated learning research (Panadero et al., Kostons et al.).
Results
- Writing quality gain: Combined group highest (M = 7.92), then FRAC (6.10), APCA (4.56), regular AI (3.10). FRAC was more closely associated with writing improvement than APCA.
- Self-assessment accuracy: APCA group had the lowest absolute error (M = 2.25) and lowest overconfidence rate (5.0%); baseline-adjusted models confirmed APCA's clearest effect, while FRAC did not significantly improve SAA.
- Effective adoption rate (EAR): Combined (0.75) > FRAC (0.62) > APCA (0.58) > regular AI (0.45); p < 0.001.
- Revision depth: Deep revision (L3+L4) rose from 28% (regular AI) to 32% (APCA), 45% (FRAC), 48% (combined).
- Transfer & retention: FRAC and combined groups kept their advantage on a new topic (T2) and after AI support was removed (T3).
Connection to Existing Knowledge Base
- AI Feedback Quality: Provides causal evidence that feedback-processing training raises effective adoption and deep revision, not just feedback volume.
- Feedback Loop: FRAC's Filter–Reason–Act–Check sequence operationalizes a structured feedback–revision loop.
- Self-Regulated Learning / Metacognition: APCA's calibration cycle is a concrete calibration-training instantiation tied to SRL monitoring and control.
- Formative Assessment: Links feedback uptake and self-assessment accuracy to formative writing assessment in higher education.
- Writing / AI Literacy: Evidence that AI-assisted writing outcomes depend on training learners to evaluate and act on AI feedback.
Methodological Notes
Strengths include a controlled factorial design with process data (decision sheets, calibration logs), multiple outcome dimensions, and baseline-adjusted self-assessment accuracy models. Limitations acknowledged by the author: a sample confined to English majors at one university, the APCA group's lower raw self-assessment error at baseline, and EAR measured directly only in the FRAC and combined groups.
What this means for practice
- Instructors. Train students in a feedback-processing routine — Filter, Reason, Act, Check — rather than simply increasing the volume of AI feedback, because that training carried the writing-quality and deep-revision gains.
- Instructors. Add a calibration cycle (Anchor, Predict, Compare, Adjust) when the target is accurate self-judgment: calibration training and feedback-literacy training improved different outcomes, and neither alone delivered both.
- Instructors. Combine the two when writing quality is the priority: the FRAC plus APCA condition produced the largest writing-quality gain and held its advantage after AI support was withdrawn.
- Learners. Treat AI feedback as a prompt to revise and re-evaluate the text rather than a verdict to accept; the regular-AI condition showed the lowest effective adoption (0.45) and the shallowest revision.
Limitations
- 120 undergraduate English majors at a single university, studied across one six-week course with four writing time points, so transfer beyond this discipline, institution, and duration is untested.
- Self-assessment accuracy was not balanced at baseline: the APCA group already showed lower raw self-assessment error, which is why the inferential SAA comparisons rely on baseline-adjusted models.
- Effective adoption rate was measured directly only in the FRAC and combined conditions, so uptake was not compared on equal terms across all four groups.
Citation
Dai, Z. (2026). Rethinking AI-assisted writing instruction: feedback literacy scripts, calibration training, and student writing development. Frontiers in Psychology, 17:1829268.