On this page

Synthesis: At the request of one biology department, Chan and colleagues coded the syllabi of every offering of its four core required biology courses — 38 syllabi in a single 2025–2026 term, a 100% collection rate — and surveyed the department's instructors to ask which graded categories they consider vulnerable to academic dishonesty. The result is the most concrete statement in the knowledge base of how much of a grade is exposed: instructors perceived only in-person proctored exams as minimally vulnerable, and on average about a third of a student's grade (33%) was highly vulnerable and roughly 80% (81%) was at least somewhat vulnerable to academic dishonesty. The uncomfortable detail for proctoring is that outside-of-class exams were rated as similarly vulnerable whether they were proctored or not, and the uncomfortable detail for evidence-based assessment reform is that the outside-of-class low-stakes assignments the field has spent thirty years promoting are exactly the category instructors regard as most exposed. The authors frame this as a grading-design problem that threatens the validity of course grades rather than a student-morality problem, and argue the responsibility belongs to departments and institutions rather than individual instructors.

Context: Two Changes That Compound

The paper's argument rests on a collision of two shifts. Thirty years ago a biology grade rested on a few high-stakes, in-person proctored exams that made dishonesty difficult and risky. Following national calls to replace high-stakes testing with multiple low-stakes formative assessments (AAAS's Vision and Change, Bransford et al., 2000), grades were rebuilt around frequent assessment, and early concerns that this would invite cheating were largely set aside: students were thought to engage authentically, plagiarism was checkable with tools like Turnitin, and outside-of-class cheating through search, collusion, or outsourcing was labor-intensive, costly, or dependent on social connections — while cheating on low-stakes work would undermine the student's own performance on later summative exams. The authors call this the pre-GenAI assumption, and their case study is a test of whether it survives.

The first change is the pandemic-era normalization of remote and asynchronous delivery, which shifted assessments outside the physical classroom and left some instructors unwilling to return to in-person proctoring even after face-to-face teaching resumed, alongside a durable increase in fully online enrolments. The second is ubiquitous access to generative AI, which raises the speed and ease of obtaining accurate answers to almost anything an instructor expects to be answered unaided, while remaining difficult to prevent or detect — driving down the perceived risk of the dishonesty that previously functioning deterrents relied on. Neither change alone exhausts the problem; the paper's claim is about their additive effect on assessments that are both outside the classroom and unproctored.

Design: Syllabi Plus Instructor Perception

The study was conducted at the request of a department at a large research-intensive US institution that offers fully online and in-person biology degrees treated as equivalent. Students in both routes take the same four core courses: introductory biology I and II and upper-level biology I and II, delivered asynchronously online for online degrees and synchronously in person for in-person degrees. The institution's size means multiple offerings of each course each term, taught by different instructors.

Two data sources were combined. First, the team collected a syllabus for every offering of all four courses in one semester of 2025–2026 (38 syllabi, 100% collection rate) and used deductive content analysis to code the graded categories each syllabus contained and the percentage of course points each category carried. The codebook defined categories on three criteria drawn from prior dishonesty research — completed inside or outside class, proctored or unproctored, and high-stakes summative or low-stakes formative — producing five categories: outside-of-class assignments, in-class assignments, unproctored outside-of-class exams, proctored outside-of-class exams, and proctored in-person exams. One graduate student and three undergraduates coded independently and met to resolve disagreements by consensus. Courses with a lab component (15–40% of the grade) were excluded because syllabi did not permit the assignment breakdown to be recovered; the lecture component was scaled to 100%.

Second, the department chair sent a survey to all instructors who had taught any biology course in the department that year — a deliberate choice of a broader population than the syllabus set, both to obtain a workable sample given typical faculty response rates and so instructors could answer honestly knowing the invitation went to over 100 people and responses were anonymous. 56 instructors responded (47% response rate). Because most teach several courses, each rated the largest course they taught, and each rated only the graded categories their own course used, to avoid judgments about categories they no longer teach. The item was "To what extent do you perceive that the following categories in your course are vulnerable to academic dishonesty?" on a scale from minimally vulnerable (0) to highly vulnerable (4). Three instructors reviewed the items for face validity before distribution. Differences across categories were tested with a Kruskal–Wallis test (χ²(4) = 67.4, p < .001) followed by pairwise Wilcoxon rank-sum tests with Bonferroni correction.

How the Grade Is Weighted

The syllabi show how differently the two modalities are built, and the differences are large. Among the 21 in-person course syllabi, 100% included outside-of-class assignments, worth a mean of 26.6% (SD 9.1) of the grade, and 90.5% (19) included in-class assignments at 14.6% (SD 6.9). All in-person courses had exams: 52.4% (11) used in-person exams at a mean of 66.5% (SD 14.8), 38.1% (8) used proctored outside-of-class exams at 54.2% (SD 2.4), and 9.5% (2) used unproctored outside-of-class exams at a mean of 50% (SD 14.1).

Among the 17 online syllabi, every course contained outside-of-class assignments, worth a mean of 40.8% (SD 7.8), and every course contained proctored outside-of-class exams, worth 59.2% (SD 7.8). No online course used unproctored outside-of-class exams. More than half the syllabi overall (61%) came from courses enrolling 200 or more students, and 45% were fully online — a distribution that makes the aggregate vulnerability figure weighted toward large courses.

What Instructors Consider Vulnerable

Instructor ratings were aggregated across modalities on the reasoning that within each graded category the three dishonesty-relevant criteria do not differ between in-person and online courses. The pattern is stark and the medians separate cleanly:

  • Outside-of-class assignments: median 4 (highly vulnerable) — the single most exposed category.
  • Proctored outside-of-class exams: median 3 (moderately vulnerable) and unproctored outside-of-class exams: median 3 — rated the same.
  • In-class assignments: median 2 (somewhat vulnerable), the category with the widest disagreement, with responses spanning all five levels.
  • In-person exams: median 0 (minimally vulnerable) — the only category instructors regarded as essentially protected.

Pairwise comparisons found in-person exams significantly less vulnerable than every other graded category (p_adj < .01) and outside-of-class assignments significantly more vulnerable than in-class assignments (p_adj < .01). The paper is careful about the interpretation of the in-person result: instructors do not regard in-person exams as immune, but their responses reflect the perceived effectiveness of in-person proctoring relative to everything administered outside the classroom.

The proctoring result is the one with the largest practical implications and the one that complicates the standard remedy. Instructors rated proctored and unproctored outside-of-class exams as equally vulnerable — that is, the responses show no perceived benefit from the available proctoring tools (lockdown browsers and similar) for exams taken outside class. The authors place this against prior literature that is genuinely mixed, with some studies reporting online proctoring as largely ineffective and others reporting reductions in dishonesty, and they note that the variation could reflect differences in both the technology and its implementation. They also note the documented costs of those tools — elevated student anxiety and concerns about privacy, technical problems, and false accusations — and draw the consequence for practice: instructors in in-person courses should weigh carefully whether the logistical savings of outside-of-class exams justify the integrity exposure, and universities should assess whether current outside-of-class proctoring technology works at all, since the alternative for online courses is assessing high-stakes learning some other way.

How Much of a Grade Is Exposed

Combining the coded point weights with the instructor vulnerability ratings produces the study's headline numbers. Treating the category instructors perceive as highly vulnerable — outside-of-class assignments — as the conservative estimate, an average of 33% of course points were highly vulnerable to academic dishonesty: 27% in in-person courses and 41% in online courses. Treating the four categories rated significantly more vulnerable than in-person exams (outside-of-class assignments, in-class assignments, and both proctored and unproctored outside-of-class exams) as the less conservative estimate, an average of 81% of course points were at least somewhat vulnerable — about 65% in in-person courses and 100% in online courses. The authors stress what this does and does not mean: it is not a claim that students are cheating, but that if a student were tempted, the assessed work representing this share of the grade could be completed dishonestly.

Because instructors retain autonomy over grading, offerings of the same course differ widely. The paper's illustration uses two in-person sections of one course: Class 1 had 10% of its points highly vulnerable, Class 16 had 33%. A student who earns a C on in-person exams while cheating to an A on all highly vulnerable assessments would finish with a C in Class 1 but a B in Class 16; if that student instead cheated on everything at least somewhat vulnerable, the grade would remain a C in Class 1 but become an A in Class 16. The same exam performance therefore certifies different things in different sections of the same course — a comparability failure as much as an integrity one, and one that raises the possibility that students select sections on the basis of how cheatable or grade-friendly they are perceived to be.

Key Findings

  1. Only in-person proctored exams were perceived as minimally vulnerable (median 0). Every other graded category was at least somewhat vulnerable, and the differences across categories were significant (χ²(4) = 67.4, p < .001).
  2. Outside-of-class assignments were perceived as highly vulnerable (median 4). They were also significantly more vulnerable than in-class assignments (p_adj < .01) and, unlike in-person exams, made up 100% of both in-person and online course syllabi.
  3. Proctoring outside class made no perceived difference. Proctored and unproctored outside-of-class exams both received a median rating of 3 (moderately vulnerable).
  4. Roughly a third of the grade was highly vulnerable. On average 33% of course points (27% in in-person courses, 41% in online courses) came from a category the instructors who use it call highly vulnerable.
  5. Roughly 80% of the grade was at least somewhat vulnerable. Averaging across the four categories rated more vulnerable than in-person exams produced 81% overall — about 65% in in-person courses and 100% in online courses.
  6. Identical performance can earn different grades in different sections of the same course. In one comparison, cheating on highly vulnerable work moved a hypothetical student from C to B in one section but left them at C in another; cheating on all at-least-somewhat-vulnerable work produced an A in the first and a C in the second.
  7. The authors locate the problem in grading structure, not student character. They call for rethinking course structure and assessment design to preserve the validity of course grades as indicators of student learning, and for departmental rather than individual responsibility.

Recommendations and Their Trade-offs

The paper's proposals are graded by severity, and each is paired with its cost. A first step is to reduce the point weight of vulnerable low-stakes assessments and shift points toward assessments that more reliably reflect learning, such as proctored in-class work — but the authors immediately note that removing incentive can also remove the reason authentically engaged students did the work, so the reform risks trading dishonesty for lost learning among the honest. A more thorough step is to adopt ungrading, uncoupling assignments from points entirely, which the authors theorize can reduce the stress of making mistakes and support intrinsic motivation — while acknowledging that a student under time pressure or lacking motivation can still complete low-stakes work with AI and no meaningful cognitive effort, so the vulnerability survives for exactly the students most likely to exploit it. The most extreme option is to eliminate outside-of-class assignments or make them ungraded practice, which the authors flag as the clearest case of the trade-off: it removes the incentive for dishonest completion but also the incentive for honest effort, and the low-stakes work still has genuine learning value for students who engage with it as intended. Their preferred direction retains incentives while enhancing proctoring and accountability, which they note is easier in person and in smaller classes — larger courses can add graduate teaching assistants or learning assistants circulating during in-class work, and in-class assignments can be structured so the question does not appear in the polling system, geolocation is enabled, and physical presence is cross-checked against a collected worksheet.

The authors also ask why instructors continue to use categories they themselves call highly vulnerable, and offer three explanations that are structural rather than personal. Longstanding disciplinary and institutional norms produce inertia, especially where pedagogical change is not incentivized or teaching identity is weak. Many instructors recognize the problem but lack time or knowledge to address it — Multimodal AI, more accurate generative AI and free institutional access both arrived in September 2024, so instructors here have had roughly a year, and research on how to design courses around these capabilities is thin. And some resist changes that conflict with prior best practice: deterring dishonesty by redesigning away from graded outside-of-class work may mean giving up flipped classrooms and active learning structures that instructors adopted because they improved learning. The paper closes by describing the department's own response — a retreat to raise awareness, a task force to develop policy guidance, and the explicit intention that the responsibility to respond should not fall on individual instructors — and argues that while one department cannot establish generalizability, the ubiquity of generative AI and the similarity of assessment architectures across universities make it unlikely that this is a single institution's problem.

What the Design Can and Cannot Establish

The study's limits follow from its design, and the paper states the central one itself: it is a case study of one department at one institution in one term, so it cannot speak to generalizability — a prediction, not a finding, is offered that the problem is widespread. Three further constraints are worth recording. The vulnerability figures rest on instructor perceptions rather than observed dishonesty, so they measure the size of the opportunity rather than its use; the study cannot say how many students act on it, and the authors are explicit that it makes no such claim. The syllabus analysis covers only lecture components, excluding lab work that constituted 15–40% of the grade in some courses, because syllabi did not disclose the assignment breakdown — so the percentages describe the graded portion that could be coded, not the whole course. And the sample is weighted toward large courses (61% enrolling 200+) and online sections (45%), which is also where the vulnerability concentrates, so the aggregate sits closer to the exposure of the largest offerings than to a typical small seminar.

What the design does establish is a method that is cheap, replicable, and diagnostic, and a finding that reframes an institutional conversation. The paper's most useful contribution to practice may be the procedure: code every syllabus's graded categories by inside/outside class, proctored/unproctored, and stakes, then multiply by instructor-perceived vulnerability to obtain a per-course exposure figure. That calculation is available to any department without new data collection, and it converts an abstract worry about AI and cheating into a concrete number that a curriculum committee can act on — which is also why the result matters beyond biology: the same grading architecture, and the same reform history, describes most of undergraduate STEM. The comparison the paper most invites is with summative redesign strategies at the other end of the spectrum, where departments have moved the certified weight of a course onto supervised, process-visible work and treated take-home assignments as practice rather than evidence (AI in Particle Physics Education: Research Problems and Foundational Skills), and with the computing-education finding that redesign, not detection, is where the evidence actually sits (Generative AI in computing education: A systematic review and a framework for responsible integration).

Connected Concepts

  • Academic Integrity — the case study's subject: how much of a grade could be earned dishonestly
  • Assessment — graded categories, point weights, and the structure of a course grade
  • Assessment Validity — grades as a warrant about student learning, undermined by unverifiable authorship
  • Formative Assessment — the low-stakes, outside-of-class assessments the field promoted are the most exposed category
  • Summative Assessment — supervised high-stakes exams as the only category instructors rated minimally vulnerable
  • Remote Proctoring — proctored and unproctored outside-of-class exams rated equally vulnerable
  • Biology Education — core required biology courses across online and in-person degree routes
  • Higher Education — institutional norms, instructor autonomy, and department-level responsibility
  • Generative AI — the technological change that removes the cost and risk of outside-of-class dishonesty
  • Curriculum Design — rebalancing points, ungrading, and the trade-off with active-learning structures
  • STEM Education — why the grading architecture described here is not biology-specific

Connected Articles

Citation

Chan, B. G., Anderson, E. P., Lu, S., Abdellatif, N., Cooper, K. M., & Brownell, S. E. (2026). Can students cheat their way to a biology degree? A case study of the vulnerability of biology course grades to academic dishonesty in the era of generative AI. OSF preprint.

Embed this page

Copy the code below to embed a chromeless version of this page in a learning management system or other website. The embedded view hides the site header, navigation, and footer.