Research Article
AI-Powered Feedback System for Writing: A Review of Automated Writing Evaluation (AWE) Tools
Synthesis: This review synthesizes 50 empirical studies published between 2019 and 2024 on how AI-powered feedback shapes Writing. The authors find that Automated Writing Evaluation (AWE) tools such as Grammarly, Quillbot and DeepL Write improve writing precision, fluency and grammatical proficiency through immediate, customized Feedback, and that they support organization and Self-Directed Learning. Three themes organize the evidence: writing accuracy improvement (25 studies), engagement and motivation (20 studies) and reduced teacher workload (15 studies). It also documents risks — weaker Critical Thinking and Creativity in learners who lean heavily on AWE, shallow discourse-level feedback, and tools tuned to standardized English norms. It positions AWE as supplementary Scaffolding: Formative Assessment works best when automated Automated Assessment feedback is paired with teacher mediation, peer review and explicit Feedback Literacy instruction.
Key Findings
- Across 50 studies from 2019 to 2024, AWE tools improved writing accuracy; 25 studies reported gains in grammar, punctuation and sentence structure from immediate Feedback.
- Accuracy improvement was the largest theme at 41.7%, ahead of student engagement and motivation at 33.3% (20 studies) and teacher workload and interaction at 25% (15 studies).
- Personalized feedback raised engagement, prompting repeated revision and greater Self-Directed Learning as students tracked their own progress between drafts.
- AWE tools cut teachers' mechanical correction workload, freeing time for higher-order concerns — argument development, coherence, organization — and for more teacher-student interaction.
- Overreliance on Automated Assessment feedback was linked to weaker Critical Thinking and creativity, while tools still misidentify correct structures and give thin discourse-level guidance.
What AWE tools change about writing accuracy
The strongest evidence concerns accuracy. Twenty-five studies reported that Automated Assessment tools improved grammar, punctuation and sentence construction, and that learners using Grammarly and Quillbot reduced grammatical mistakes and produced more complex sentences across successive drafts. The mechanism is iterative: because Feedback arrives immediately and can be requested again and again, students revise many times before submission rather than waiting for one teacher response. Accuracy improvement accounted for 41.7% of findings, the largest share. AWE also sits alongside Automated Essay Scoring in the broader family of automated Assessment in Language Learning. Several studies found gains in vocabulary range and coherence, and one six-month study recorded measurable growth in overall writing proficiency — but only where tools were embedded in structured drafting, feedback and revision cycles rather than used as stand-alone aids.
Motivation, independence and the risk of overreliance
Twenty studies, 33.3% of findings, linked AI-powered feedback to stronger engagement and motivation. Learners valued feedback that was personalized, available at any time, and more comprehensive and reliable than expected; that confidence encouraged more detailed revision and ownership of their progress. AWE supported Self-Directed Learning: students could monitor improvement and spot gaps without waiting for a teacher. The same evidence warns that learners depending too heavily on automated suggestions became less critical and less willing to make substantive revisions, drifting toward the system's preferences rather than a developing academic voice — a pattern close to Cognitive Offloading. Because perceptions are subjective, with some students finding the feedback motivating and others boring or excessively critical, the motivational benefit is not guaranteed. The authors tie that variability to individual attitudes and to Feedback Literacy.
The changing role of the teacher
Fifteen studies, 25% of findings, documented a reduction in teachers' corrective workload once AWE tools entered the writing classroom. Automated correction of mechanical errors let instructors spend less time reading every paper for surface mistakes and more on higher-order issues such as argument development, topic selection, coherence and organization. That reallocation let teachers interact more with students, discuss ideas and analytical thinking, and give the tailored guidance automated systems cannot provide. The review does not read this as replacement. Human Feedback remains essential for motivational support, for argument structure and rhetorical impact, and for helping students interpret harsh comments. The authors return repeatedly to teacher mediation as the condition that makes Formative Assessment with AI pedagogically coherent, noting that early teacher feedback keeps later Self-Directed Learning productive in large or remote Higher Education classes.
Cultural and contextual limits of automated feedback
Several limits recur across the studies. AWE tools often misidentify correct structures as errors or miss complex syntactic patterns, eroding learner trust. Most emphasize form and offer little help with discourse-level concerns such as audience awareness, coherence and rhetorical strategy. A deeper problem is cultural: tools built on standardized English norms may not represent the rhetorical traditions of learners from other linguistic backgrounds, risking the marginalization of legitimate styles, particularly in EFL settings. Students using Grammarly reported dissatisfaction with evaluations that followed American spelling and punctuation conventions. Digital literacy also shapes how learners engage with AWE regardless of second-language level, and AI cannot fully capture the diversity of written English. The authors therefore call for pairing automated feedback with human guidance, peer review and collaborative learning so accuracy is not pursued at the expense of voice.
What this means for practice
- Instructors. Treat AWE output as a first-pass scaffold: build drafting cycles around the tool, then discuss its suggestions so students learn to judge when Feedback is wrong.
- Instructors. Spend your commenting time on higher-order work — argument, coherence, audience and voice — where the reviewed studies found automated systems weakest.
- Instructors. Teach Feedback Literacy directly: have students compare AWE comments with peer review and with your own, and name the styles the tool undervalues.
- Program leaders. Embed AWE in a structured Writing sequence instead of offering it as a stand-alone tool, and fund the teacher mediation learners still need.
Limitations
- The review covers only empirical English-language peer-reviewed studies published from 2019 to 2024 in Google Scholar, ERIC and ResearchGate, excluding theoretical work and non-ELT contexts.
- Study counts are inconsistent: the abstract and method describe 50 studies selected for coding, while the findings section discusses 60, and the theme percentages are computed on that larger base.
- Tool coverage centers on Grammarly, Quillbot and DeepL Write with no links or versions given, and effectiveness rests largely on self-reported perceptions of the tools.
Citation
Gres, E., Riyanti, D., & Yuliana, Y. G. S. (2026). AI-Powered Feedback System for Writing: A Review of Automated Writing Evaluation (AWE) Tools. Wiralodra English Journal, 10(2).