AI Ed Wiki logoAI Ed WikiUse with AI

Synthesis: Farrokhnia et al. (2026) run a randomized three-group experiment with 70 university students to compare teacher feedback against ChatGPT feedback produced with two prompting techniques — Zero-shot and chain-of-thought (CoT) — for argumentative essay writing in Persian. They find that CoT prompting yields significantly higher-quality feedback than both Zero-shot prompting and a human teacher, but that this quality advantage does not translate into greater essay revision gains. Teacher feedback, rated lower in quality, produces comparable improvement in revised essays. The authors conclude that feedback quality alone is insufficient; students' engagement with and uptake of feedback are the decisive factors, motivating a hybrid model in which teachers help students interpret and apply GenAI feedback.

Core Finding

Higher-quality AI feedback does not automatically produce better student revisions — what matters is whether and how students engage with and act on the feedback. In a randomized experiment, chain-of-thought prompting produced objectively higher-quality feedback on argumentative essays than both Zero-shot prompting and an experienced human teacher, yet the CoT group did not revise its essays significantly more than the teacher-feedback group, which saw comparable gains. This decoupling of feedback quality from learning gains is the paper's central and somewhat counterintuitive result: it challenges the assumption that improving the "quality" of generated feedback is sufficient to improve writing outcomes, and redirects attention to students' active feedback uptake and interpretation.

Prompt Engineering and Feedback Quality

The study directly interrogates Prompt Engineering as a determinant of GenAI feedback quality. The Zero-shot prompt gave ChatGPT a direct instruction to generate feedback from an argumentation rubric, while the CoT prompt guided the model through a step-by-step evaluation with an elaborated example. One-way ANOVA (F(2,67) = 6.09, p = .004, Ī·p² = 0.15) showed CoT feedback (M=12.90) significantly outperformed both Zero-shot (M=11.25, p=.01) and teacher feedback (M=11.20, p=.008), with no significant difference between Zero-shot and teacher. The authors interpret this as CoT's stepwise reasoning aligning GenAI outputs more closely with the cognitive demands of argumentative writing — and frame prompt design within explainable AI principles. This is a valuable empirical contribution to the wiki's Prompt Engineering and AI Feedback Quality concepts.

Why Quality Did Not Translate into Revision Gains

The finding that teacher feedback — rated lower in quality — produced comparable revision improvements highlights the critical role of Feedback Literacy and student Agency. The authors note that high-quality feedback should be specific and actionable, but its effect depends on students' willingness and ability to implement it. Notably, GenAI feedback quality was significantly associated with students' initial essay quality, whereas teacher feedback quality showed no such association — meaning GenAI responded differently depending on how strong the initial draft was, while the teacher calibrated more consistently. The study's Persian-language setting also extends GenAI-feedback research beyond English-dominant contexts, testing generalizability in a linguistically underrepresented language.

Implications for Practice

The authors advocate for hybrid intelligent feedback systems in which teachers scaffold students' interpretation and application of GenAI feedback, rather than treating AI as a standalone replacement for the instructor. This aligns the paper with the wiki's Human AI Collaboration and Teacher Role concepts, and with Writing Education practice: GenAI can generate rich, structured, scalable feedback, but the human teacher remains essential for helping students engage with it meaningfully. For Assessment and Formative Assessment, the result cautions against assuming better AI feedback automatically yields better learning.

Relevance to the Wiki

This is a tightly controlled experimental contribution to the wiki's feedback cluster. It provides causal, comparative evidence that links AI Feedback Quality, Prompt Engineering, and learning outcomes in Higher Ed, and it resonates strongly with the wiki's existing coverage of AI-generated feedback, essay scoring, and teacher-vs-AI comparisons. It also gives concrete guidance for Instructional Design: prompt technique matters for feedback quality, but Pedagogy (scaffolding uptake) matters for learning.

Connected Concepts

Connected Articles

Citation

Farrokhnia, M., Latifi, S., Papadopoulos, P. M., Hogenkamp, L., Gijlers, H., Khosravi, H., & Noroozi, O. (2026). Generative AI offers more, but students revise less: comparing the effects of teacher and AI feedback on student essay revisions. International Journal of Educational Technology in Higher Education.