Research Article
Student Evaluation of Repeated AI Feedback Across a Semester of Writing
Synthesis: This short paper provides rare descriptive classroom evidence on what happens when students repeatedly use generative-AI feedback across a full semester of writing coursework. Drawing on 2,988 reflective essay-feedback-appraisal instances from 283 Estonian bachelor students, the authors find that students rated AI feedback as helpful and actionable more often than not, but a growing minority (about one in ten) found it unhelpful toward the end of the term. The work sits squarely in the Artificial intelligence and feedback in university education: effectiveness and student perceptions literature and complements prior AI Feedback Quality studies by tracking feedback appraisal longitudinally rather than in a one-off lab task.
The study surfaces the central tension in Over-Reliance: generative AI offers a fast, scalable route to immediate writing advice, but it is not a self-contained path to deeper reflection. Using a validated AI-text classifier, the authors estimate the share of essays that look like unaided student writing, linking tool use to the broader question of whether AI assistance erodes learning gains. These findings reinforce concerns echoed in Faster Completion, Less Learning: Generative AI Reduced Study Time on Math Problems and the Knowledge They Build and Assessing the Impact and Underlying Pathways of Sequenced AI Feedback on Student Learning about dosage and critical engagement. The paper argues benefits depend on whether students learn to use AI selectively and critically, a skill squarely within AI Literacy and the Student Experience of writing support in Writing.
What this means for practice
- Learners. Evaluate the feedback instead of accepting it: appraisals shifted over the semester from general acceptance to identifying generic advice, mismatch with the essay and misreadings, so judge each comment for relevance, trustworthiness and usability.
- Learners. Give the tool the essay context that makes advice specific: prompt-level grounding reduced generic-advice complaints but did not settle judgments about accuracy, tone or actionability, so test advice against your own text.
- Learners. Record which feedback you accept or reject and why: the study captured cohort-level appraisals, not whether individual students became better evaluators, and the completion-oriented rubric made full credit easy to obtain without revision.
- Learners. Do the reflective work yourself: descriptive reflection depth did not change across the semester, and roughly a third of essays were classified as likely generated writing - including 51% of those coded as dialogic or critical reflection.
Limitations
- One course and one assignment format: 2,988 appraisal instances from 283 Estonian bachelor students, and the authors state they cannot determine whether growing negativity and AI overuse came from general semester fatigue rather than the task or course.
- Students chose their own chatbots and provider labels were derived from URLs rather than exact model settings, so the specific tool behind each appraisal is uncontrolled.
- Appraisals were a short feedback-appraisal slot within reflective essays and were analyzed with LLM-assisted coding, which remains an approximation of students' judgment.
- The AI-text classifier used to estimate essay groundedness does not officially support Estonian and was assessed on a small test set, and the authors note a large share of the corpus likely cannot be treated as purely unaided student writing.
Citation
Karjus, A., Leoste, J., & Oun, T. (2026). Student Evaluation of Repeated AI Feedback Across a Semester of Writing.