Research Article
Peer and AI Review + Reflection (PAIRR): A Human-Centered Approach to Formative Assessment
Synthesis: Sperber et al. (2025) present the Peer and AI Review + Reflection (PAIRR) model, a human-centered approach to formative assessment that combines peer review best practices with AI review while emphasizing student agency and reflection. In the largest study of college students' use of AI feedback to date (N = 654 across 10 writing courses and three writing-intensive STEM courses at UC Davis), they found that AI feedback is most useful when combined with peer review: the majority of students preferred combined feedback, found the similarity between AI and peer feedback reassuring, valued their complementary perspectives, and developed AI literacy by critically assessing AI outputs.
Key Findings
- The majority of students (58%) preferred to receive feedback from both ChatGPT and peers, while 36% preferred peer feedback alone and only 6% preferred AI feedback alone — AI utility was experienced in the context of human feedback.
- AI and peer feedback were often similar and mutually reinforcing: 75% of students reported similarities between peer and ChatGPT feedback, which they described as confirming and strengthening each other, increasing confidence in peer response.
- When AI and peer feedback differed, they were complementary: AI feedback was often described as "overly general" (31%) but provided actionable revision strategies and rubric-driven feedback on organization, focus, and structure, while peer feedback was more specific and detailed (28%), drew on contextual knowledge of the course and assignment, and provided an authentic audience and emotional support.
- By evaluating AI outputs rather than taking them at face value, students developed AI Literacy — including learning ethical ways to use AI — while many asserted writerly agency in deciding whether to accept or reject feedback (only 5.3% of students demonstrated overconfidence in AI feedback).
- Continued feedback conversations increased perceived utility: 35% of students continued a ChatGPT conversation, and 71% of those preferred combined feedback (vs. 50% who did not). Self-Efficacy did not predict continuation, and preferences did not differ statistically between small writing classes and large WI courses.
Study Design & Method
This mixed-methods study implemented PAIRR in 10 distinct writing courses plus three large writing-intensive (WI) courses at a large R1 public university in the western US during winter and spring 2024. The 654 participating students were diverse (37% first-generation, 13% international, 68% multilingual). The intervention sequence: students read and reflected on articles about AI; drafted, provided/received peer review, then prompted ChatGPT for rubric-driven feedback; critically assessed both kinds of feedback and made revision plans; and revised and reflected on the process. Data included pre-/post-surveys, interviews with 4 faculty and 11 students, focus groups with 8 TAs and 1 reader, and 654 students' reflections and feedback assessments. Survey data were analyzed with descriptive/inferential statistics (binary logistic regression and chi-squared test in R); qualitative data were thematically coded in MaxQDA using a stratified sub-sample of 131 students (20%) with open, axial, and selective coding and inter-rater reliability.
What this means for practice
- Instructors. Sequence feedback as peer review, then rubric-driven AI feedback, then critical comparison and a revision plan: of 654 students, 58% preferred combined feedback, against 36% preferring peers alone and 6% preferring AI alone.
- Instructors. Have students judge both sources side by side rather than choosing one, since 75% saw similarities that reinforced confidence while the differences proved complementary — AI feedback was "overly general" to 31% and peer feedback more specific for 28%.
- Faculty developers. Make reflection and revision planning a required step, because assessing AI output is where students built AI literacy and asserted writerly agency (only 5.3% showed overconfidence in AI feedback).
- Instructors. Encourage students to continue the AI conversation when it is productive: 35% did so, and 71% of those preferred combined feedback versus 50% of those who did not.
- Administrators. Scale the model to large writing-intensive courses as well as small classes: preferences did not differ statistically by course size across the 10 writing courses and three WI courses studied.
Limitations
The study's focus was on student perceptions of AI feedback utility, so it did not directly evaluate AI outputs for bias or quality. Differences in course mode, writing support, instructor experience, rubrics, assignment prompts, and genre across courses may have affected perceptions; the large online PLA course was overrepresented, and its students received only one peer reviewer (vs. two elsewhere) yet got comprehensive TA feedback. ChatGPT 4.0 was released mid-study, and version access was variable. The authors did not distinguish one-shot vs. continued-conversation feedback in the main findings.
Citation
Sperber, L., MacArthur, M., Minnillo, S., Stillman, N., & Whithaus, C. (2025). Peer and AI review + reflection (PAIRR): A human-centered approach to formative assessment.