On this page

Luo and Dawson (2026) ask the question students keep asking: will I get lower grades if I use GenAI? Their answer is that there may never be a definitive one. Drawing on scenario-based interviews with 33 university teachers across the Greater Bay Area of China, the study shows that grading GenAI-assisted work is not a technical application of criteria but an exercise in value judgment — teachers conjecture about who the student is, what they are capable of, how they relate to others and whether the decision does good.

The sample skews toward junior and teaching-track staff (25 of 33 had 0–5 years' experience; 16 were teaching assistants, post-doctoral fellows, language instructors or adjunct teaching fellows), which the authors justify because these teachers spend more time grading. None of the participating universities had policies on grading GenAI-assisted work or had banned student GenAI use, leaving teachers substantial autonomy.

Synthesis: An exploratory qualitative interview study of 33 teachers documenting how grading GenAI-assisted work becomes a site of competing, largely unexamined values. Four clusters emerge — person-oriented (honesty, diligence), capability-oriented (independence, GenAI skill, disciplinary mastery), relation-oriented (trust) and justice-oriented (fairness, beneficence) — across 66 value-judgment instances. The authors explain the variation via epistemological stages, disciplinary traditions and sociocultural context, then argue that validity should govern marking down and that the remedy is two-way transparency.

Key Findings

  1. Grading is a value judgment, not neutral measurement: teachers prioritize often-implicit values, and many here were unaware of the orientations behind their own grading.
  2. Person-oriented values (22 instances) turn on honesty and diligence: concealed use was read as dishonesty, use as laziness, but failure to engage with GenAI as a lack of diligence.
  3. Capability-oriented values (33 instances) were most frequent, but what counts as capability became murky: independence emerged as the key marker, implicitly ranking "completely human" work above AI-assisted work.
  4. The same teacher can hold competing values: one industrial design teaching assistant graded on acknowledged quality, then instinctively ranked the AI-free student higher, then conceded equal grades were fairest.
  5. Relation-oriented values (7 instances) center on trust — built over time, and threatened when honest disclosure is penalized despite policies permitting GenAI use.
  6. Justice-oriented values (4 instances) cover fairness (resentment among independent students, incentives to game the system) and beneficence (avoiding harm from false accusation).
  7. Validity is the proposed pathway forward, paired with two-way transparency about whether GenAI use is acceptable and how it will affect grades.

Four value clusters, one messy space

Person- (22) and capability-oriented (33) values were invoked far more often than relation- (7) and justice-oriented (4) ones. Ethics cuts across all of them, since fairness, honesty and trust are morally loaded, though the authors make ethics no standalone cluster to avoid overlap; they also note that perceived diligence could tip borderline cases. The practical result is inconsistency: a teacher may reward skillful GenAI use as evidence of critical thinking in one scenario and penalize declared use in the next, while perceived diligence could tip borderline cases. The authors call this a "messy grading space full of tension and inconsistency," warning that leaving it unaddressed invites student distrust, biased grading and weakened credibility of academic certification.

Explaining the variation

The study lacks data to conclude why teachers differ, so the authors offer probable explanations at three levels. Individually, Kuhn, Cheney and Weinstock's (2000) epistemological stages apply: absolutists treat GenAI-assisted work as deviating from prescribed demonstrations of knowledge; multiplists, seeing all approaches as equally valid, may let effort or honesty drive the grade; evaluativists weigh which methods are better supported by evidence. Disciplinarily, humanities teachers were more critical of GenAI use, while hard and applied sciences teachers reported fewer challenges — a nursing lecturer's assignments were face-to-face role plays, a civil engineering teaching assistant's were projects and in-person exams. Socioculturally, an implicit "AI shaming" hierarchy favoring human-only effort operated against teachers' stated openness; the authors speculate that a Confucian context prizing hard work and honesty explains the heavier weighting of person- and capability-oriented values.

Validity and two-way transparency

Validity — the degree to which grades represent what they are meant to — is the organizing remedy. Under a validity view, GenAI use justifies marking down work if and only if it meant students could not demonstrate the outcomes being assessed, the position of English writing teachers who believed GenAI had interfered with students showing their writing skills. Otherwise the authors find substantial construct-irrelevant variance, where grades vary on factors unrelated to the assessed outcomes: marking down because GenAI use suggests a student is not diligent, or declining to mark down while GenAI did the substantive work, both threaten validity. Since students and assessors are unclear where the line lies, the mitigation is transparency in both directions.

What this means for practice

  • Instructors. Examine the value judgments behind your grades — honesty, effort, trust — and decide whether your course rewards skillful GenAI use or independent work; left implicit, this costs grades their validity.
  • Faculty development. Workshops should build evaluativist reasoning: whether critical thinking must be demonstrated unaided or with GenAI.
  • Assessment design. Review rubrics against what the assessment is meant to assess; at minimum, state your stance on GenAI in assignments and its grading consequences.
  • Program leaders. Where teaching assistants carry major grading responsibility (16 of 33 participants), instructors and graders need clear communication to grade consistently.
  • Researchers. Triangulate — observe real grading with think-aloud protocols, and analyze policies, rubrics and work samples.

Limitations

  • The scenario-based method means responses may not reflect actual grading behavior; research shows teachers can espouse progressive practices while practicing conservatively.
  • The scenarios were built from the researchers' own judgments about what is controversial, so framing effects may have shaped answers.
  • The sample sits in one national context, the Greater Bay Area (six universities), and skews junior: 2 associate professors or professors and 25 of 33 with 0–5 years' experience.
  • No participating university had GenAI grading policies or a ban, so findings describe grading under substantial teacher autonomy rather than a defined policy regime.

Connected Concepts

Connected Articles

Citation

Luo, J. (Jess), & Dawson, P. (2026). Exploring value judgements in grading: will teachers mark down student work assisted by GenAI, and should they?. Studies in Higher Education, 51(9), 1970–1984.

Embed this page

Copy the code below to embed a chromeless version of this page in a learning management system or other website. The embedded view hides the site header, navigation, and footer.