On this page

Synthesis: This paper proposes a principled framework grounded in Evidence-Centered Design (ECD) that treats Generative AI as a design variable within STEM assessment arguments rather than an external threat. This represents a significant evolution beyond the binary debate of 'ban AI vs. allow AI' that has dominated discussions about Academic Integrity in education.

Three Governance Stances

The framework articulates three context-dependent AI Governance stances based on how GenAI interacts with the assessment's validity argument:

  1. Restrict — warranted when GenAI would contaminate the inferential chain between student work products and targeted unaided proficiency. This preserves the validity of assessments designed to measure independent competence.
  2. Scaffold — warranted when bounded GenAI support can assist with peripheral demands without revealing the target construct, preserving inferential interpretability. This aligns with Scaffolding approaches in Intelligent Tutoring systems.
  3. Require — warranted when the target construct is disciplinary AI interaction competency itself. Tasks elicit process artifacts (prompts, critiques, revisions) that make student reasoning observable and scorable, distinguishing it from AI-generated output.

Empirical Validation

Two task designs deployed in an introductory physics course demonstrated that disciplinary AI interaction competencies are observable in student response artifacts and can be scored using defensible rubrics grounded in student data and expert knowledge. This connects to Automated Grading and Confidence Estimation in Automatic Short Answer Grading with LLMs research on making student reasoning visible and scorable.

What this means for practice

  • Instructors. Set the AI-use rule per assessment by asking how GenAI would reshape the inference from a student's work product to the targeted construct, instead of applying a blanket allow-or-ban policy.
  • Instructors. When AI interaction is the target construct, require process artifacts — prompts, output critiques and assumption reflections — so each student's own reasoning stays separable from the AI output.
  • Administrators. Ground assessment AI decisions in the validity argument rather than institutional preference, and keep Educational AI Policy guidance and AI Literacy standards tied to that argument so learning integrity is preserved while students are prepared for AI-enabled workplaces.
  • Designers. Confirm an accuracy ceiling before assigning a Require stance: the model must consistently miss at least one expert-anchored error, otherwise uncritical acceptance of AI output is undetectable and the task collapses into ordinary AI collaboration.

Limitations

  • The framework is a theoretical analysis illustrated by two task designs in a single introductory physics course; it reports no student performance sample, and the governance stances are not empirically tested.
  • The authors state that the effectiveness of Scaffold governance is an unverified assumption: whether guardrailed AI support leaves disciplinary reasoning substantively equivalent to unassisted reasoning remains an open question.
  • Accuracy-ceiling validation rests on an expert answer key for the two tasks, not on a comparison between governed and ungoverned conditions or a broader set of disciplinary tasks.
  • The scaffold/interlocutor boundary is expected to shift as frontier models improve; the paper identifies empirical monitoring of when a task crosses a governance boundary as an unresolved methodological challenge.

Citation

Gao, Y., Chen, Z., Li, M., & Zhai, X. (2026). Generative AI as a design variable: An evidence-centered framework for principled governance in STEM assessment.

Embed this page

Copy the code below to embed a chromeless version of this page in a learning management system or other website. The embedded view hides the site header, navigation, and footer.