Liu, Song, Gallagher, Sterman & August (2026) โ University of Illinois Urbana-Champaign / Montana State University. ๐ Full text (arXiv)
Introduces FOXGLOVE, a dataset of 696 feedback comments by trained writing instructors on 69 twelfth-grade argumentative essays, paired with 1,644 comments from four frontier LLMs โ totaling 2,340 comments with expert quality ratings. Provides the first systematic comparison of LLM and expert feedback on three pedagogically critical dimensions: goal-orientation, anchoring to specific sentences, and prioritization.
Key findings: Instructors and LLMs distribute feedback similarly across revision goals and essay positions, but diverge significantly on which specific sentences receive feedback. Models write more complex feedback and use fewer questions than human instructors. LLM feedback receives higher quality ratings on most dimensions โ but much of this advantage is attributable to lengthier comments inflating perceived quality.
This work directly informs the design of AI writing feedback systems, highlighting the need to evaluate feedback quality beyond surface-level ratings and to consider pedagogical factors like feedback anchoring and prioritization. Relevant to both secondary and higher education writing instruction.
Related Pages
- writing-education โ AI-supported writing instruction and feedback
- ai-feedback-quality โ Quality evaluation frameworks for AI-generated feedback
- formative-assessment โ Formative assessment in writing through AI
- feedback-loop โ Feedback loops in AI-assisted writing
- k-12 โ Secondary education writing contexts
- higher-ed โ Higher education writing instruction