On this page

Synthesis: > AI assistance does not only make required work faster or more accurate — it can increase the amount of optional but beneficial work that actually gets done. In a semester-long randomized field experiment across a 300-level machine learning course, teaching assistants shown an AI-generated feedback draft after grading were significantly more likely to provide feedback (+10.81 percentage points) and produced longer comments (+39.79 characters), without spending more time per character or degrading students' usefulness ratings. Drafts acted as editable scaffolds that lowered the barrier to initiating feedback rather than eliminating the effort of producing it, and TAs stayed fully in control — able to use, edit, or ignore every draft. The study reframes AI's role from a productivity tool to an intervention that changes whether discretionary, socially valuable work happens at all.

Key Findings

  1. AI-assisted feedback drafts increased feedback provision by +10.81 percentage points (SE = 1.10, p < 0.001) in a setting where giving feedback was optional, not required.
  2. Feedback length grew by +39.79 characters (SE = 3.45, p < 0.001); because submissions without feedback were coded as zero, this estimate captures both more provision and more content per student.
  3. Time spent per character of feedback did not change significantly (0.29 s/char, SE = 0.35, p = 0.41), suggesting AI supported task initiation rather than removing the work of reviewing, adapting, and deciding whether to send feedback.
  4. Students rated AI-assisted and non-assisted feedback as equally useful (difference −0.01, SE = 0.06, p = 0.88), so more feedback did not degrade perceived quality or Student Experience.
  5. Qualitative interviews showed TAs treated drafts as editable intermediate artifacts that helped them articulate and verify comments they already had in mind — not as final outputs to be trusted wholesale.

Context: Feedback as Discretionary Work

A large body of AI-assisted-workflow research evaluates systems on settings where the task is required — engineers completing assigned coding tasks, physicians producing diagnoses. But many socially valuable practices, such as mentoring, documentation, reviewing, and feedback provision, are discretionary: people intend to do them yet often skip them amid competing demands. Personalized feedback in Higher Education is a canonical example. It is pedagogically valuable, improving student Motivation, learning, and experience, but it is labor-intensive and hard to scale, so TAs provide it selectively rather than uniformly.

This paper's core contribution is shifting the question from how well AI helps with a required task to whether AI makes an optional task happen at all. It situates feedback provision in a broader CSCW tradition of invisible and discretionary work, arguing that participation gaps can sometimes be closed by changing how work is surfaced, measured, or supported.

The Intervention: AI Drafts in the Grading Workflow

The study ran in a 300-level undergraduate machine learning course at a private R1 university (enrollment 130–150), where feedback was historically rare because it was optional. Written work was manually graded by 11 TAs (7 GTAs, 4 UTAs) using instructor-provided rubrics, and each question was graded by a single TA.

The intervention was a lightweight, Large Language Models (LLMs)-backed Chrome extension that integrated into the course's native grading platform. It ran a two-stage design intended to limit over-reliance: TAs first completed grading independently, and only after clicking a "Done Grading" button was an AI-assisted feedback draft surfaced. In the treatment condition the draft appeared; in the control condition it did not. TAs could use, edit, or ignore each draft at their discretion. Question-level randomization with a rotating assignment scheme across homework (HW1–HW4) meant each TA graded both treated and control submissions, with homework-question fixed effects absorbing the assigned TA, rubric, and question characteristics.

The drafts were generated by o4-mini, selected via a formative study in which five TAs ranked feedback from four candidate LLMs and a from-scratch option using a Bradley-Terry preference model. Formative results also shaped prompt design: instructors iterated on a feedback prompt that took the problem description, instructor solution, grading rubric, and student submission as inputs, and TAs preferred feedback that was specific, actionable, concise, and stylistically direct.

Findings

RQ1 — Feedback Provision Behavior

The behavioral results were consistent with the two headline effects: a +10.81pp increase in feedback provision and a +39.79-character increase in feedback length, both highly significant and stable across homework assignments. Critically, time per character did not change (0.29 s/char, p = 0.41). Because the feedback-length estimate codes no-feedback submissions as zero, it bundles the "whether" and the "how much" effects together. The flat time-per-character result implies AI primarily lowers the barrier to starting feedback, while TAs still expend real effort reviewing, verifying, and adapting each draft — the mechanism the interviews clarify.

RQ2 — Usage Patterns

Interviews with all 11 TAs revealed that drafts functioned as editable scaffolds. TAs frequently struggled to articulate feedback even when they could identify an issue; the personalized, specific drafts helped translate implicit judgments into clear comments and made it easier to phrase feedback they already wanted to give (e.g., T1: "I wouldn't know how to give them the feedback the right way... and the templates would do it in such a good way"). Drafts also supported verification — helping TAs confirm their interpretation of a student response and refine tone or specificity. TAs therefore treated the AI output as an intermediate artifact to be shaped, not as an authoritative final product, reflecting a human-in-the-loop relationship rather than delegation to an autonomous agent.

RQ3 — Student Experience

The increased provision did not degrade downstream outcomes. Students rated AI-assisted and non-assisted feedback as similarly useful (difference −0.01, p = 0.88), and interviews with a subset of 9 students suggested they valued the increased availability of personalized feedback while generally being unable to distinguish AI-assisted from non-assisted feedback. In other words, more feedback arrived, and its perceived quality was preserved.

What this means for practice

  • Instructors. Target the AI at discretionary work rather than required work: with nothing else changed, showing TAs a draft lifted feedback provision by 10.81 percentage points (SE = 1.10, p < 0.001) in a course where giving feedback was optional.
  • Instructors. Surface the draft only after grading is finished, so the AI supports articulating and verifying a judgment the TA has already made instead of substituting for it — the two-stage design that limited over-reliance.
  • Instructors. Keep drafts use-or-edit-or-ignore: the flat time-per-character result (0.29 s/char) shows the design buys task initiation, not the removal of review effort, and TAs should keep spending that effort.
  • Instructors. Deploy the pattern where personalized Feedback is pedagogically important but TA time is scarce — large-enrollment courses — and judge it on participation and perceived usefulness, not on efficiency alone.
  • Instructors. Treat human oversight as a fixed constraint, not an efficiency to optimize away; the authors are explicit that preserving review caps the possible effort savings relative to an autonomous pipeline.

Limitations

  • The randomized field experiment ran in a single 300-level machine learning course (enrollment 130–150) at one private R1 university across four homework assignments in one semester, with 11 TAs.
  • The claim that AI supports initiation rather than reducing effort is inferred from a behavioral proxy: time per character was not a direct measure of initiation cost, cognitive effort, or the work of checking and adapting drafts.
  • Several outcomes are self-reports — TA ratings of draft usefulness and student ratings of the feedback they received — and the student interviews covered a subset of 9 students.
  • Longer-run effects are untested: the authors call for work on whether gains persist across semesters, whether subject domain moderates the effect, and how sustained exposure shapes TA reliance on drafts.

Citation

Romina Mahinpei, Victoria Dean, Ruth Fong, Lydia T. Liu, Manoel Horta Ribeiro (2026). AI Assistance for Discretionary Work: Increasing Feedback Provision in Higher Education. arXiv.

Embed this page

Copy the code below to embed a chromeless version of this page in a learning management system or other website. The embedded view hides the site header, navigation, and footer.