On this page

Synthesis: Across five U.S. middle and high schools (N=143 students, grades 7–11), rubric-grounded RAG feedback from CyberScholar supported students' writing revision, with most students reporting improvements in organization, elaboration, and style.

Key Findings

  1. Students valued the detailed, criterion-specific feedback and the tool's interactive, iterative qualities, which fostered revision and reduced reliance on teacher feedback.
  2. Automated star ratings were inconsistent — some students at one site received different scores for the same unchanged submission — and occasionally misaligned with assignment expectations, underscoring the need for human oversight.
  3. Teachers reported that the tool saved time on feedback and supported more targeted, higher-order instructional practices, while some worried that overly specific suggestions could limit students' critical thinking.

Synthesis

CyberScholar demonstrates rubric-grounded RAG (Retrieval-Augmented Generation) for formative writing feedback at scale across five US schools. The tool integrates teacher-provided rubrics, materials, and exemplars through RAG to produce criterion-specific feedback — a design that directly addresses the Formative Assessment challenge of providing timely, rubric-aligned feedback without overburdening teachers. The 143 students (grades 7-11) valued the immediate, iterative feedback and reported improvements in organization, elaboration, and style. However, automated rating inconsistencies and occasional rubric misalignment highlight the continuing need for human oversight — a finding consistent with the Human-in-the-Loop principle that AI feedback should augment rather than replace teacher judgment. The teacher time-saving benefit (freeing educators for higher-order instruction) connects to Educational Development and the Teaching evolution identified in Modeling AI-TPACK in Practice: Insights from Teachers'' Multi-Agent Workflow Design. CyberScholar's rubric-grounded design also contrasts with more open-ended LLM feedback approaches studied in The Effects of Structured LLM-Generated Feedback on Programming Assignment Performance, suggesting domain-specific rubric integration as a promising direction for educational AI feedback systems.

Background

Writing is an essential 21st-century skill tied to college and employment readiness, yet it remains a challenging competence for students to acquire — only 27% of U.S. 12th-graders demonstrated writing proficiency on the 2022 National Assessment of Educational Progress. Access to detailed Feedback is constrained by teacher time, large class sizes, and workload, which has motivated interest in Generative AI as a complement to human instruction. Prior tools such as Grammarly and QuillBot, grounded in earlier Educational NLP, focused mainly on grammar and mechanics, leaving open how AI might deliver in-depth, criterion-aligned feedback during writing. This study argues that aligning GenAI feedback to teacher rubrics and instructional materials can make feedback more transparent, consistent, and actionable in K-12 settings, a claim supported by research showing rubrics clarify expectations for students and guide objective evaluation for teachers.

The CyberScholar Tool

CyberScholar is a Multimodal AI writing workspace in which students draft in a vertically split editor while AI and human dialogue tools occupy the right-hand panel. The platform integrates teacher-provided rubrics, materials, and exemplars through RAG (Retrieval-Augmented Generation), retrieving from a curated, educator-validated Knowledge Base stored in a vector database to produce criterion-specific formative feedback and ratings aligned to teacher expectations. Several distinct functions support the writing process: CyberHelper lets writers request AI assistance while drafting; CyberReviewer delivers quantitative ratings with qualitative narrative justification from AI rubric agents, peers, teachers, or self-reflection; and the Composition Report analyzes whether AI use facilitates Cognitive Offloading or extends student thinking by combining generative AI with logfile data such as keystrokes and clickstreams. Rubric Agents, built by teachers or instructional designers, align evaluation with disciplinary frameworks through multiple-pass Prompt Engineering and chain-of-thought processes. Teachers may select from open-weight or commercial foundation models, while learner identities and work remain securely separated through an application programming interface, with restrictions on data retention and persistent model training on student submissions as deliberate privacy and equity decisions.

Design and Methods

The study adopted a qualitative, interpretivist Qualitative Research approach using a multiple-case design, with each of five U.S. school sites (four high schools and one middle school) serving as a bounded case across urban, suburban, and rural settings. Participants included 143 students in grades 7–11 and five teachers; data collection combined classroom observations, student post-surveys (n=79), student focus group interviews (n=18), and teacher surveys (n=5). Analysis followed two cycles of inductive coding in which three researchers built a collaborative codebook, triangulated themes across instruments, and reached agreement on disagreements. Teachers were conceptualized as co-implementers within a Human-in-the-Loop cyber-social Pedagogies and Teaching Strategies, with onboarding emphasizing rubric agent construction, interpretation of AI feedback, calibration of ratings, and strategies to prevent overreliance — a design that foregrounds teacher professional learning and responsible GenAI implementation as integral to the intervention.

Student Perspectives

The thematic analysis surfaced four patterns in how students engaged with CyberScholar. First, students consistently praised the provision of detailed and specific feedback, noting that the tool identified concrete revision targets — exact sentences, grammar, sentence structure, and organization — and offered suggestions for improvement, which supported a productive revision process. Second, students recognized writing improvement through AI feedback, describing how criterion-based recommendations helped them strengthen conclusions, improve coherence, and refine word choice, contributing to a sense of progress and confidence. Third, students engaged with the delivery of ratings connected to rubric criteria, appreciating that the tool made abstract rubric language visible and referenced categories like "writing mechanics," though opinions were divided on the star-rating display. Fourth, where the interactive feature was available (Schools C, D, and E), students treated the tool as a conversational partner, asking follow-up questions, requesting elaboration or synonyms, and using the back-and-forth to drive an iterative cycle of feedback, revision, and reassessment that enhanced their sense of agency during revision.

Teacher Perspectives

Teachers reported that CyberScholar saved time on repetitive feedback, freeing them to focus on higher-order concerns such as argument development, coherence, and reasoning. Several noted the value of having GenAI feedback aligned to their own rubrics and instructional requests, and one emphasized that the star display worked as motivation because students treated stars as feedback rather than grades. However, teachers also voiced cautions: some worried that highly specific suggestions could limit students' opportunities for Critical Thinking and independent intellectual discovery, and others contrasted the tailored AI feedback with the contextual personalization only a teacher could provide. Teacher 02 at School B found the 4-star scale too simplistic and potentially distracting, suggesting points or percentages instead, and noted that high performers might accept a 3-of-4 rating without reading the underlying feedback. Teachers also observed that the tool's impact depended heavily on students' willingness to revise, echoing broader concerns about engagement and self-regulation.

What this means for practice

  • Instructors. Anchor the tool to your own rubrics, materials, and exemplars before students touch it. Students across all five sites valued criterion-specific feedback — exact sentences, grammar, sentence structure, organization — and reported that it told them how to revise rather than only what was wrong.
  • Instructors. Treat automated star ratings as provisional drafts, not results. Some students at one site received different scores for the same unchanged submission, and participants questioned rating reliability precisely because scores were inconsistent across submissions — calibrate before scores reach students or families.
  • Instructors. Set explicit norms for AI dialogue that protect critical thinking. Teachers worried that highly specific suggestions could limit independent discovery, and one observed that high performers might accept a 3-of-4 rating without reading the underlying feedback.
  • Instructors. Redirect the time the tool saves on repetitive mechanics into higher-order instruction: teachers reported it freed them for argument development, coherence, and reasoning, which is the trade the study actually documents.
  • Designers. Separate feedback from scoring in the learner view. One teacher found the 4-star scale too simplistic and potentially distracting and suggested points or percentages, while another valued the star display as motivation because students treated stars as feedback rather than grades — the two readings cannot be served by one display.

Limitations

  • As a qualitative multiple-case practice study focused on feasibility and participant perceptions, it cannot determine whether CyberScholar caused measurable improvements in writing achievement, whether gains were sustained, or how AI-supported revision compares with teacher-only or peer feedback in a controlled design.
  • The sample is small and uneven across sites: 143 students were involved, but only 79 completed the post-surveys and 18 took part in focus groups, with participation voluntary, so self-selection is possible.
  • Implementation spanned a single academic semester (2024/2025), which limits generalizability and raises the possibility of novelty effects.
  • The evidence is primarily self-reported perceptions, observational field notes, and thematic coding rather than independent, standardized measures of writing quality or inter-rater reliability statistics, so the study cannot establish the reliability, validity, or fairness of the automated ratings across demographic groups.

Citation

Zheldibayeva, R., de Oliveira Nascimento, A. K., Castro, V., Cope, B., & Kalantzis, M. (2026). Generative AI Feedback, English Writing and Teacher Rubrics: A Multiple-Case Study of CyberScholar.

Embed this page

Copy the code below to embed a chromeless version of this page in a learning management system or other website. The embedded view hides the site header, navigation, and footer.