On this page

Synthesis: Zhao and Ji compare three versions of the same argumentative essay — the student's first draft, the revision made after peer feedback, and a revision produced by ChatGPT-4 from that same first draft — using Appraisal Theory, which maps the evaluative language a text uses to position itself against other voices. Across 168 essays written by 14 undergraduate English majors in China, the AI revisions carried the most appraisal resources overall, mostly through Attitude and Graduation, while the frequency of Engagement resources stayed flat across all three draft types. What separated the conditions was not how much engagement language appeared but how it was configured: peer-feedback revisions became more contractive (more Counter and Endorse, less Entertain), whereas the AI revisions kept a more balanced Contract/Expand mix. The authors read this non-competitively — the two routes display different stance-taking strategies, so the AI text's pedagogical value is as a comparison object that makes dialogic positioning visible.

Key Findings

  1. The corpus was 14 undergraduate English majors (11 female, 3 male; average age 19) at one Chinese university, each writing four argumentative tasks over nine weeks — a first draft, a revision after anonymous peer feedback (two reviewers given, two received per task), and an AI-generated revision — for 168 essays: 56 of each draft type, containing 58,854 tokens and 2,883 sentence units.
  2. The AI revisions were produced by ChatGPT-4 from each student's own first draft using one standardized three-step prompt (criteria for organization, development, cohesion and coherence, sentence structure, vocabulary and mechanics; then the writing prompt; then the draft), so all three conditions shared one instructional context and criteria set.
  3. Total appraisal resources rose across the conditions — first draft lowest, peer-feedback revision higher, AI-generated revision highest — driven by Attitude (estimated marginal means 11.574 → 15.677 → 18.429) and Graduation (4.737 → 5.074 → 6.006), not by Engagement (8.142 → 7.849 → 7.999), which did not differ significantly by draft type (F(2, 326) = 0.380, p = .684).
  4. Within Engagement, Heteroglossic resources dominated Monoglossic ones in every condition (M = 6.140 vs 1.871, F(1, 326) = 725.525, p < .001), and that split stayed stable across draft types (interaction F(2, 326) = 0.522, p = .594).
  5. The Contract/Expand balance did shift by condition (interaction F(2, 314) = 5.681, p = .004): peer-feedback revisions became more contractive (Contract M = 3.602 against Expand M = 2.380, compared with 3.222 and 2.970 in first drafts), while AI-generated revisions were relatively more balanced (Contract M = 3.456, Expand M = 2.789).
  6. At the subcategory level, Counter increased in both revision conditions (0.840 → 1.241 → 1.338), Endorse rose sharply in peer-feedback revisions (0.131 → 0.384) but fell in AI-generated ones (0.100), Pronounce declined in both (0.547 → 0.461 → 0.361), and Entertain was the most frequent resource throughout (2.551 → 2.001 → 2.367).
  7. The authors explain the AI pattern through documented chatbot tendencies — hedging and expansive resources from human-feedback alignment, alongside contrastive devices such as "however" and "while" — and note that the balanced mix resembles what the literature associates with more proficient argumentation; the same balance did not appear in a comparison study of essays generated from scratch rather than revised from a student draft.

Reading revision as dialogic positioning

Appraisal Theory treats a text as a set of evaluative moves rather than a container of claims. Its Engagement subsystem is the part that matters for argument: Monoglossic resources assert without acknowledging other voices, while Heteroglossic resources bring other voices in — Expand (Entertain, Acknowledge, Distance) opens dialogic space, and Contract (Negation and Counter, plus the Proclaim moves Concur, Endorse, Pronounce and Justify) closes it. Earlier work in this tradition found L2 writers, and Chinese EFL writers in particular, favoring Monoglossic and contractive choices while L1 English writers expand more — a difference in stance-taking, not in grammatical accuracy. The distinction matters beyond this setting because writing instruction usually evaluates revision by whether the text got better (clearer thesis, better evidence, fewer errors) rather than by what the revised text now does rhetorically — the same dimension AI feedback research leaves implicit when it asks whether feedback improved a draft. A model can improve a paragraph's prose while making the writer's stance more monologic, and no rubric would show it.

What each revision condition produced

The conditions differed in configuration, not volume: Engagement frequency was statistically indistinguishable across the three draft types, and so was the Monoglossic/Heteroglossic split. Both revision conditions used Contract more than Expand overall (M = 3.427 vs 2.713), but peer-feedback revisions pushed further in the contractive direction — Counter and Endorse rose and Entertain fell, so the text closed more dialogic space than the draft it came from. The authors' interpretation is that students defending their own claims, and appealing to external authority to legitimize them, may be treating assertiveness as a proxy for a strong argument, a tendency reinforced when peers evaluate the work through the same rhetorical expectations the class was taught.

The AI revisions moved differently. Their Counter use was the highest of the three conditions (1.338), yet Entertain stayed comparatively high and Pronounce dropped furthest, so the text combined a stronger counter-claiming move with continued hedging. Two examples show the mechanism: one revision opens an opposing view with "Admittedly … some may … argue" before returning to its own position (Concede + Entertain + Distance), and another concedes that peer feedback is "well-intentioned" before arguing that it "may fail to catch subtle language errors" (Counter + Concede + Entertain). The authors connect this to how chat assistants are trained to hedge and to their habitual contrastive connectives, while flagging that the pattern did not appear in a study of AI essays written from scratch.

Neither distribution is intrinsically superior: contractive positioning can be exactly what a claim needs, and expansive positioning can thin out an argument. What the comparison supplies is a diagnostic — the two routes make different rhetorical choices visible, which matters because students rarely see the alternatives to their own default stance. It also reframes the AI revision as a contrast text to argue with rather than a model answer to copy, which is why the authors position it as complementary to peer feedback rather than a substitute for it.

What this means for practice

  • Teach stance explicitly, not just structure. The authors' suggestion is to put two sentences side by side — one that Entertains ("This may imply that …") and one that Counters ("On the contrary, this viewpoint ignores …") — and ask students what each does to the space for alternative views, introducing appraisal terminology through the comparison rather than as a vocabulary lesson.
  • Run peer feedback and AI revision as a sequence, not a competition. Have students revise from peer feedback first, then compare their revision with an AI revision of the same draft and identify where the two diverge in handling opposing views. Guiding questions: does the expression strengthen or soften the claim, how does it position alternatives, and which version serves the writer's purpose?
  • Use the AI text to break the assertiveness-equals-strength equation. Students here moved toward closing down dialogue after peer feedback, which the authors read as a plausible misconception about what makes argument persuasive. Comparing revisions is a low-stakes way to show that conceding a point and countering it are both available moves.
  • Judge revisions on positioning as well as quality. A checklist built only on organization, evidence and language can miss whether the writer's stance became more or less open — a dimension worth naming in Feedback to student writers and in the criteria instructors use to evaluate AI-generated feedback on argumentative work.

Limitations

  • The sample is 14 undergraduate English majors at one Chinese university, so the patterns describe one EFL instructional context rather than a general result.
  • All AI revisions came from a single standardized prompt, so the observed Contract/Expand balance may partly reflect that prompt and its criteria rather than a general property of AI revision.
  • The AI revisions came from ChatGPT-4, a superseded generation of the assistants a reader can open today: the appraisal profile describes that model's revision behavior under this prompt, and the paper's own comparison with an earlier model-generation study already showed the pattern is not fixed.

Citation

Zhao, H., & Ji, J. (2026). Human-written and AI-generated revisions: An appraisal analysis of dialogic positioning in argumentative writing. Frontiers in Education, 11, 1904852.

Embed this page

Copy the code below to embed a chromeless version of this page in a learning management system or other website. The embedded view hides the site header, navigation, and footer.