๐Ÿง  AI Ed Wiki

Synthesis: Unravelling undergraduates' development of evaluative judgments through AI-supported internal feedback

Key Findings

  • This qualitative case study followed 28 second-year pre-service teachers (18 female, 10 male; Mage = 20.6) in a compulsory English writing course at a university in northern China across a 14-week semester with three writing tasks (150โ€“250-word essays at College English Test Band-6 level).
  • The intervention used DeepSeek R1 to generate strong, average, and weak exemplars; students ranked their own draft against the three exemplars, justified the rankings on self-reflection forms, then sought AI feedback, compared it against the criteria and their goals, and revised.
  • Through iterative analogical comparison (draft vs. exemplars) and analytical comparison (draft vs. AI feedback, criteria, and goals), students' evaluation focus expanded from language features to higher-order aspects of writing โ€” content, organisation, and coherence โ€” although their evaluative reasoning generally lacked sophistication: example-based justification doubled from 15% to 30.5% across tasks while simple statements declined, but the most sophisticated comparative reasoning remained rare.
  • Internal feedback shifted over the semester from reflections on writing skills and criteria (54.5% and 33.3% of codes in Task 1) toward metacognition, goal alignment, and self-monitoring โ€” with metacognitive reflections rising exponentially from 3% in Task 1 to 24.3% in Task 3.
  • Three evaluator types emerged with distinct developmental pathways: reconstructive (N = 8), language-focused (N = 9), and criteria-compliance (N = 11) evaluators.
  • Study Design & Method

    The study is grounded in the internal feedback paradigm: students generate feedback by comparing their work against external information (exemplars, criteria, goals, AI feedback) โ€” with the comparison itself producing the learning gain. Each task cycle had students write an initial draft, set goals, rank their draft against three AI-generated exemplars of varied quality and justify the ranking, request and compare DeepSeek R1 feedback, revise, and set goals for the next task. The instructor (a senior lecturer with over 15 years of experience) explained the writing criteria in advance. Data sources were students' self-reflection forms, semi-structured interviews, and initial and revised essay drafts, analysed qualitatively to trace evaluative judgment development over time.

    Key Results

  • Expanding evaluation focus: students' attention moved from language-level features toward content, organisation, and coherence across the three tasks, indicating growth in what they considered quality writing.
  • Reasoning quality: most students could discern features of quality writing but mainly indicated presence or absence of features without deeper explanation. Example-based reasoning rose (15% โ†’ 30.5%), simple statements fell steadily, but comparative analysis โ€” the most sophisticated form โ€” did not emerge.
  • Internalization trend: attention to writing skills and criteria dropped (54.5% โ†’ 21.6% and 33.3% โ†’ 5.4% respectively), while metacognition, goal alignment, and self-monitoring rose โ€” consistent with the strategies becoming embedded in writing practice and iterative goal-setting cycles catalyzing self-monitoring.
  • Evaluator types: reconstructive evaluators expanded their focus from language to higher-order dimensions but remained descriptive in reasoning; language-focused evaluators (like "Phoebe") consistently prioritized grammatical accuracy and vocabulary, refining rather than broadening their conception of quality; criteria-compliance evaluators (N = 11) anchored judgments to the stated criteria. AI feedback shaped each pathway differently โ€” e.g., one language-focused student doubted DeepSeek's Band-6 advice until the instructor confirmed it, refining her understanding of exam requirements.
  • Implications for AI in Education

    The study demonstrates a practical strategy for turning GenAI into a scaffold for assessment literacy rather than a shortcut: comparing drafts against AI-generated exemplars of varied quality helps students appreciate a "quality continuum," while comparing drafts with AI feedback and personal goals makes the internal feedback explicit and auditable. Having students articulate their internal feedback on self-reflection forms both reduces cognitive load and lets teachers track progress. The three developmental pathways imply that one-size-fits-all prompting fails โ€” teachers should customize self-reflection prompts to each student's evaluative orientation (e.g., pushing language-focused evaluators toward organisation and coherence, while reinforcing evidence-based reasoning for reconstructive evaluators). This connects to Self Regulated Learning and Feedback Loop research: students become agentic seekers of AI feedback who monitor goal attainment and judge the contextual appropriateness of AI suggestions, echoing concerns in Assessment about students' critical evaluation of AI-generated feedback.

    Limitations

  • The accuracy of students' rankings and evaluative reasoning was not examined (resource constraints), and the course instructor's perspective was not elicited during data analysis despite member-checking with some students.
  • Participants came from a single institution and discipline, limiting transferability across disciplinary and institutional settings.
  • Students sought external feedback mainly from AI, and not all AI feedback is accurate โ€” which may have influenced evaluative judgment development. The authors recommend teachers comment on AI feedback quality and guide students to evaluate it critically.
  • Connected Concepts

  • Self Regulated Learning
  • Assessment
  • Feedback Loop
  • Metacognition
  • AI Tutoring
  • Reducing AI Misuse
  • Math Education
  • Prompt Engineering
  • Connected Articles

  • Bloom Aligned Educational Control Llms โ€” From Execution to Education: A Bloom-Aligned Framework for Measuring Educational Control in LLMs
  • Multimodal Learning GenAI โ€” Multimodal Learning with Generative AI
  • AI Learning Assistants Higher Ed Large Scale โ€” Using AI-based Learning Assistants in Higher Education: A Large-Scale Descriptive Analysis
  • LLM Automated Assessment Student Self Explanations โ€” Exploring the Effectiveness of Using LLMs for Automated Assessment of Student Self Explanations in Programming Education
  • Youtube Frames Chatgpt Education โ€” How YouTube Frames ChatGPT Use in Education: An Epistemic Network Analysis with Supporting Multimodal Metadata
  • GenAI Performance Vs Learning โ€” Distinguishing performance gains from learning when using generative AI
  • Citation

    Chen, S., & To, J. (2026). Unravelling undergraduates' development of evaluative judgments through AI-supported internal feedback.