๐Ÿง  AI Ed Wiki

Jennifer M. Krebsbach & Victoria L. Cross (University of California, Davis) โ€” Assessment & Evaluation in Higher Education (Taylor & Francis). Open Access, CC BY 4.0. doi:10.1080/02602938.2026.2686727.

Jennifer M. Krebsbach & Victoria L. Cross (University of California, Davis) โ€” Assessment & Evaluation in Higher Education (Taylor & Francis). Open Access, CC BY 4.0. doi:10.1080/02602938.2026.2686727.

Summary

A natural-experiment / design-based study tracking eight iterations of a lower-division Data Visualisation in the Social Sciences course (n = 921 across six years) to test whether โ€” and how โ€” GenAI changes student learning. The authors frame faculty anxiety about GenAI as the latest in a series of "moral panics" (calculators, word processors, search engines, e-learning) and argue the productive response is to teach and embed GenAI use, not ban it. They compare three instructional conditions on two quiz types (knowledge vs. applied):

  • pre-GenAI (2019โ€“2020, n = 3 cohorts)
  • GenAI-available (2023โ€“2024, n = 3) โ€” GenAI present but no pedagogical adaptation; some students used it, often ineffectively/unethically
  • GenAI-integrated (2025, n = 2) โ€” explicit instruction + encouragement to use GenAI on the applied portion; GenAI banned on the knowledge portion (paper quiz)
  • Method (key design)

  • Quizzes 3โ€“6 analysed (first two dropped as orientation; Quiz 7 dropped as low-stakes). Item-level performance (% correct = difficulty; SD = variability) from the LMS.
  • 3 ร— 4 mixed-design ANOVA: AI-availability (between, 3 levels) ร— quiz number (within, 4 levels). Small per-condition N (2โ€“3 cohorts), so effect sizes (ฯ‰ยฒ) reported as the primary evidence.
  • Key Findings

    Applied questions โ€” "available" hurt, "integrated" recovered

  • Main effect of GenAI availability: F(2,5) = 5.85, p = 0.049, ฯ‰ยฒ = 0.35 (GenAI availability accounts for 35% of variance in applied-question performance).
  • In the GenAI-available condition, applied performance was significantly lower than baseline on Quizzes 4, 5, 6. Because applied questions could not be answered directly by GenAI, the drop indicates students were less prepared โ€” either unable to use GenAI effectively or unable to critically evaluate its output.
  • In the GenAI-integrated condition, applied performance returned to ~pre-GenAI levels (and exceeded baseline on one harder quiz). Teaching students to use GenAI for data summarising levelled the field.
  • Knowledge questions โ€” availability masked cheating; paper quiz revealed a deficit

  • Same main effect, ฯ‰ยฒ = 0.35. Knowledge performance stayed at baseline during GenAI-available, then dropped below baseline once delivered on paper in the integrated condition.
  • The authors interpret the stable central tendency but elevated variability during GenAI-available as evidence that some students unethically used GenAI to boost knowledge scores (heterogeneous use masked underlying learning differences). Moving knowledge quizzes to paper removed that opportunity and exposed that integrated-cohort students were less prepared โ€” plausibly from over-reliance on GenAI to summarise content.
  • Variability โ€” the headline signal

  • Applied-question variability: F(2,5) = 64.84, p < 0.001, ฯ‰ยฒ = 0.88 โ€” GenAI availability accounted for 88% of variability. Variability spiked in the GenAI-available condition, returned to baseline under integration.
  • Knowledge-question variability: availability ร— time interaction F(6,15) = 11.65, p < 0.001, ฯ‰ยฒ = 0.57; main effect ฯ‰ยฒ = 0.78. GenAI availability accounted for 78% of the increase. Paper delivery produced the lowest, most stable variability โ†’ read as greater equity in the classroom.
  • Student feedback (pilot, n = 28/149 responded)

    Mixed: 53% preferred the new split format (paper knowledge + take-home applied); common praise was reduced stress and more active calculation. Others preferred the old 25-min efficiency.

    Interpretation: design beats ban

    The integrated redesign resolved both academic-integrity and authenticity concerns by splitting the quiz: a high-integrity paper knowledge test + a high-authenticity, open-resource applied task where GenAI use was taught. The authors caution they likely over-learned the "moral panic" lesson โ€” assuming universal, effective GenAI adoption โ€” when in reality uptake was partial and often ineffective. Their conclusion: monitor our own hypotheses about student GenAI use, keep learning objectives central, and design authentic assessments for the new environment rather than condemn the technology.

    Connected Concepts

  • Plagiarism Detection
  • Reducing AI Misuse
  • Automated Essay Scoring
  • Student Experience
  • AI Misuse Learning Harm
  • AI Literacy
  • Student Misconceptions AI
  • Higher Ed
  • Connected Articles

  • Beyond Detection Authentic Assessment AI 2025 โ€” Beyond Detection: redesigning authentic assessment in an AI-mediated world
  • Authentic Products Authenticated Processes 2026 โ€” From authentic products to authenticated processes: authentic assessment in AI-rich higher education
  • Tool Invariant Framework Agentic AI โ€” A Tool-Invariant Framework for Teaching and Assessing Computational Methods in the Age of Agentic AI
  • Code Review GenAI Cs1 โ€” Combating Harms of Generative AI in CS1 with Code Review Interviews and a Flipped Classroom
  • Teaching Intro AI Course Redesign Bill Of Rights 2026 โ€” Teaching Intro AI When the Tools Can Do the Homework: A Course Redesign and a Student Bill of Rights
  • GenAI Skill Bypass Literacy โ€” The GenAI Skill Bypass: Mapping Divergent Pathways of University Students and Staff AI Literacy
  • Citation

    Krebsbach, J. M., & Cross, V. L. (2026). Navigating the moral panic: encouraging appropriate use of GenAI in the classroom rather than condemning innovation as disruption. Assessment & Evaluation in Higher Education. https://doi.org/10.1080/02602938.2026.2686727