Research Article
Positioning Generative AI in EFL Peer Feedback: Training Feedback Literacy and Enabling Uptake in Speaking Classes
Synthesis: Irwin & Muller (2026) offer a conceptual, theory-informed design framework for positioning generative AI (GenAI) within EFL/ESL peer feedback on speaking, targeting the twin bottlenecks of peer feedback production (variable, generic, or criterion-misaligned comments) and uptake (cognitive overload from interpreting multiple, sometimes contradictory peer messages). Grounded in student feedback literacy (Carless & Boud, 2018) and a cyclical, self-regulated model of GenAI-enabled feedback engagement (forethought–control–retrospect; Zhan et al., 2025), the paper proposes two complementary GenAI roles: a Trainer that supports feedback givers through exemplar-based calibration, rubric-guided practice, and feedback-on-feedback; and a Synthesizer that supports feedback receivers by aggregating peer comments into concise, prioritized, criteria-linked uptake reports. Speaking tasks are argued to impose distinct constraints — time pressure, fleeting performances, and heightened affect — that make real-time peer feedback promising yet pedagogically challenging. The proposal is deliberately scoped to performative speaking assessment in higher education, positions GenAI as scaffold rather than evaluator (teachers remain in the loop; no automated grading), and concludes with four design principles, AI Governance considerations, eight testable theoretical expectations, and a future research agenda.
Key Findings
- Two GenAI roles: The Trainer (for feedback givers) uses exemplars and analytical rubrics so learners calibrate judgments against expert ratings and revise vague comments toward criterion-referenced, evidence-based, actionable feedback (e.g., turning "Nice voice" into a time-coded intelligibility diagnosis with practice suggestions). The Synthesizer (for feedback receivers) normalizes formats, groups comments by criterion, preserves minority views, and generates a concise uptake report with prioritized themes, contradictions to clarify, candidate next-step goals, and linked learning resources.
- Cyclical mapping: The Trainer operates mainly in the forethought and early-control phases (building appreciation, calibrated judgment, and affect management), while the Synthesizer operates in later control and retrospect (supporting taking action, reflection, and revision efficacy).
- Speaking-specific constraints: Time pressure, fleeting oral performance, and heightened affect make real-time peer feedback challenging; the paper draws on evidence that asynchronous, recorded, revisitable workflows and multiple peer raters stabilize judgments and support reflection.
- Design principles: (1) feedback timing and sequencing are critical (pre-task training primes criteria; immediate post-task feedback supports engagement; spacing supports reliable ratings); (2) Trainer units should be short, repeated, and task-aligned to manage cognitive load; (3) the Synthesizer should preserve the givers' voice while simplifying intake; (4) Guardrails remain essential — teacher oversight, no automated grading, CEFR-level-matched prompts/resources, and offline fallbacks for low-infrastructure contexts.
- Ethics and governance: Framed as "generativism" (symbiotic human–AI collaboration; Pratschke, 2024), the design sends students' anonymized feedback (not oral production) to GenAI to protect Privacy, keeps prompts and rubrics teacher-designed, and retains original giver comments so outputs can be audited for hallucination or bias. Macro- vs. microethical framing (Kubanyiova, 2008) is used to focus on classroom-specific (microethical) concerns.
- Eight theoretical expectations: Trainer is expected to improve attitudes, feedback quality, judgment alignment, and feedback literacy; Synthesizer is expected to support uptake/action planning and reduce overload/affect; combined use is expected to support speaking development over time and more equitable participation. All remain provisional.
- Boundary conditions and limitations: The account is conceptual (no new empirical data); scope is limited to EFL/ESL performative speaking in higher education; assumes adequate digital infrastructure; and acknowledges Large Language Models (LLMs) output unreliability (hallucination risk) and the risk of overreliance on GenAI or crowding out human dialogue about feedback.
What this means for practice
- Instructors. Sequence peer feedback rather than assigning it: run pre-task Trainer calibration with exemplars and a rubric, scaffold comment composition through the task, then issue post-task Synthesizer uptake reports linked to criteria and learning goals.
- Instructors. Keep Trainer units short, repeated, and task-aligned so calibration does not itself overload learners, and account for speaking-specific constraints — time pressure, fleeting performances, and heightened affect — by using asynchronous, recordable workflows and more than one peer rater.
- Instructors. Keep the givers' voice and the teacher in the loop: preserve the original peer comments so Synthesizer output can be audited, leave grading with teachers, and let GenAI support rather than replace the evaluative judgment peers are meant to build.
- Designers. Protect Privacy by sending only anonymized feedback notes rather than students' oral production to GenAI, keeping prompts and rubrics teacher-designed, and retaining original comments so outputs can be checked for hallucination or bias.
- Administrators. Build the Guardrails before adopting the workflow: CEFR-level-matched prompts and resources, no automated grading, teacher-in-the-loop oversight, and an offline fallback for contexts without adequate digital infrastructure.
Limitations
- The framework is conceptual and adds no empirical data: the eight theoretical expectations are propositions to be tested, not findings.
- Scope is limited to EFL/ESL performative speaking in Higher Education; the design does not address written or product-based peer feedback.
- It assumes adequate digital infrastructure — the proposed offline fallback is a design gesture, not an evaluated component.
- The authors acknowledge that LLM output is unreliable (hallucination risk) and that students may over-rely on GenAI, or the workflow may crowd out the human dialogue about feedback that peer review is meant to produce.
Citation
Irwin, B., & Muller, T. (2026). Positioning generative AI in EFL peer feedback: Training feedback literacy and enabling uptake in speaking classes. Education Sciences, 16, 544. https://doi.org/10.3390/educsci16040544