FOXGLOVE: Comparing Goal-Oriented Writing Feedback from Experts and LLMs

Created: 2026-06-09 | Tags: llmwriting-educationformative-assessmentfeedback-loophigher-edk-12

Liu, Song, Gallagher, Sterman & August (2026) โ€” University of Illinois Urbana-Champaign / Montana State University. ๐Ÿ“„ Full text (arXiv)

Introduces FOXGLOVE, a dataset of 696 feedback comments by trained writing instructors on 69 twelfth-grade argumentative essays, paired with 1,644 comments from four frontier LLMs โ€” totaling 2,340 comments with expert quality ratings. Provides the first systematic comparison of LLM and expert feedback on three pedagogically critical dimensions: goal-orientation, anchoring to specific sentences, and prioritization.

Key findings: Instructors and LLMs distribute feedback similarly across revision goals and essay positions, but diverge significantly on which specific sentences receive feedback. Models write more complex feedback and use fewer questions than human instructors. LLM feedback receives higher quality ratings on most dimensions โ€” but much of this advantage is attributable to lengthier comments inflating perceived quality.

This work directly informs the design of AI writing feedback systems, highlighting the need to evaluate feedback quality beyond surface-level ratings and to consider pedagogical factors like feedback anchoring and prioritization. Relevant to both secondary and higher education writing instruction.

Related Pages

Citation

APA: Liu, Y., Song, Y., Gallagher, J., Sterman, S., & August, T. (2026). FOXGLOVE: Understanding Goal-Oriented and Anchored Writing Feedback from Experts and LLMs on Argumentative Essays. arXiv:2606.06271.