On this page

Synthesis: A natural-experiment / design-based study tracking eight iterations of a lower-division Data Visualization in the Social Sciences course (n = 921 across six years) to test whether — and how — GenAI changes student learning. The authors frame faculty anxiety about GenAI as the latest in a series of "moral panics" (calculators, word processors, search engines, e-learning) and argue the productive response is to teach and embed GenAI use, not ban it. They compare three instructional conditions on two quiz types (knowledge vs. applied):

  • pre-GenAI (2019–2020, n = 3 cohorts)
  • GenAI-available (2023–2024, n = 3) — GenAI present but no pedagogical adaptation; some students used it, often ineffectively/unethically
  • GenAI-integrated (2025, n = 2) — explicit instruction + encouragement to use GenAI on the applied portion; GenAI banned on the knowledge portion (paper quiz)

Method (key design)

  • Quizzes 3–6 analyzed (first two dropped as orientation; Quiz 7 dropped as low-stakes). Item-level performance (% correct = difficulty; SD = variability) from the LMS.
  • 3 × 4 mixed-design ANOVA: AI-availability (between, 3 levels) × quiz number (within, 4 levels). Small per-condition N (2–3 cohorts), so effect sizes (ω²) reported as the primary evidence.

Key Findings

Applied questions — "available" hurt, "integrated" recovered

  • Main effect of GenAI availability: F(2,5) = 5.85, p = 0.049, ω² = 0.35 (GenAI availability accounts for 35% of variance in applied-question performance).
  • In the GenAI-available condition, applied performance was significantly lower than baseline on Quizzes 4, 5, 6. Because applied questions could not be answered directly by GenAI, the drop indicates students were less prepared — either unable to use GenAI effectively or unable to critically evaluate its output.
  • In the GenAI-integrated condition, applied performance returned to ~pre-GenAI levels (and exceeded baseline on one harder quiz). Teaching students to use GenAI for data summarizing levelled the field.

Knowledge questions — availability masked cheating; paper quiz revealed a deficit

  • Same main effect, ω² = 0.35. Knowledge performance stayed at baseline during GenAI-available, then dropped below baseline once delivered on paper in the integrated condition.
  • The authors interpret the stable central tendency but elevated variability during GenAI-available as evidence that some students unethically used GenAI to boost knowledge scores (heterogeneous use masked underlying learning differences). Moving knowledge quizzes to paper removed that opportunity and exposed that integrated-cohort students were less prepared — plausibly from over-reliance on GenAI to summarize content.

Variability — the headline signal

  • Applied-question variability: F(2,5) = 64.84, p < 0.001, ω² = 0.88 — GenAI availability accounted for 88% of variability. Variability spiked in the GenAI-available condition, returned to baseline under integration.
  • Knowledge-question variability: availability × time interaction F(6,15) = 11.65, p < 0.001, ω² = 0.57; main effect ω² = 0.78. GenAI availability accounted for 78% of the increase. Paper delivery produced the lowest, most stable variability → read as greater equity in the classroom.

Student feedback (pilot, n = 28/149 responded)

Mixed: 53% preferred the new split format (paper knowledge + take-home applied); common praise was reduced stress and more active calculation. Others preferred the old 25-min efficiency.

Interpretation: design beats ban

The integrated redesign resolved both Academic Integrity and authenticity concerns by splitting the quiz: a high-integrity paper knowledge test + a high-authenticity, open-resource applied task where GenAI use was taught. The authors caution they likely over-learned the "moral panic" lesson — assuming universal, effective GenAI adoption — when in reality uptake was partial and often ineffective. Their conclusion: monitor our own hypotheses about student GenAI use, keep learning objectives central, and design authentic assessments for the new environment rather than condemn the technology.

What this means for practice

  • Instructors. Redesign the assessment rather than banning the tool: the 2025 GenAI-integrated condition recovered applied-question performance to pre-GenAI levels by moving knowledge questions to a paper quiz and teaching GenAI use on the open-resource applied portion.
  • Instructors. Teach the summarizing workflow explicitly, because in the GenAI-available condition applied performance fell significantly below baseline on Quizzes 4, 5 and 6, which the authors read as students being less prepared or unable to evaluate GenAI output.
  • Instructors. Monitor score spread, not only the mean: GenAI availability accounted for 88% of the variance in applied-question performance (omega-squared = 0.88) and 78% of knowledge-question variability (omega-squared = 0.78), and paper delivery produced the lowest, most stable variability - read as greater equity.
  • Instructors. Design authentic, open-resource applied tasks that GenAI cannot answer directly, since the authors argue the split quiz resolves Academic Integrity and Authentic Assessment concerns at once.
  • Researchers. Treat cohort patterns as hypotheses about student behavior: the authors caution that they over-learned the moral panic lesson by assuming universal and effective GenAI adoption when uptake was partial and often ineffective.

Limitations

  • Small-N design: eight course sections across three conditions (pre-GenAI n = 3, GenAI-available n = 3, GenAI-integrated n = 2), even though each quiz drew 84-202 students; the authors report effect sizes as the primary evidence and urge caution with the inferential statistics.
  • Retrospective and non-experimental: students were only indirectly exposed to different treatment conditions across six years, so causation cannot be determined and alternative explanations for the cohort differences remain open.
  • With eight sections, assumptions of normality and sphericity cannot be meaningfully evaluated, and the authors acknowledge the mixed-design ANOVA was chosen despite that limit.
  • The format pilot drew 28 respondents from 149 students (53%, n = 15, preferred the split format) and ran on the final quiz, so the authors note that the positive sentiment may be inflated.

Citation

Krebsbach, J. M., & Cross, V. L. (2026). Navigating the moral panic: encouraging appropriate use of GenAI in the classroom rather than condemning innovation as disruption. Assessment & Evaluation in Higher Education.

Embed this page

Copy the code below to embed a chromeless version of this page in a learning management system or other website. The embedded view hides the site header, navigation, and footer.