Concept
Educational AI Policy
Educational AI policy — the formal and informal rules governing AI use in educational institutions, from national legislation to classroom guidelines. Policy research in the knowledge base spans institutional governance, curriculum mandates, and teacher preparation requirements.
Questions to Consider
- Institutional AI policies are often written as if telling people the rules changes what they do. Research suggests these policies can lag far behind actual AI use. Why do you think a policy 'on paper' so often fails to match what happens in classrooms?
- If homework outsourcing causes learning losses that go largely unnoticed — because no single teacher connects a student's decline to AI use — what kind of policy would even be possible? What would monitoring look like?
- Assessment policy decisions like oral exams, proctored tests, and closed-book formats are themselves responses to AI-enabled cheating. Do you think policing the 'output' (catching AI use) is a better strategy than redesigning assessment so the work itself is harder to outsource?
- Educational AI policy spans national legislation, government guidance, institutional rules, and classroom-level choices. Which of these levels do you think actually shapes student and teacher behavior the most — and why?
- Whose voices tend to shape AI policy — and whose are missing? Who should be at the table when an institution decides what AI use is allowed and how it's governed?
Introduction
Policy levels
- Institutional policy: Institutional policy analysis compares how universities develop AI policies. Institutional change frameworks provide models for policy development. Qian (2026) coded the principal AI pages of the 50 US universities ranked most innovative in 2025 and found that only 6 framed them as a "policy" while 44 published guidance, guidelines, principles or resource hubs, and that 33 of 50 positioned instructors as the primary translators of institutional expectations into course rules — institutional policy is being deliberately written as adaptable guidance for course-level interpretation rather than fixed rule.
- Government policy: and lifelong learning policy examine regulatory approaches at national and regional levels. A critical, infra-level view of AI in government policy-making is Perrotta (2026), whose analysis of the UK Redbox civil-service Large Language Models (LLMs) prototype shows how general-purpose AI enters the professional toolkit of policy through "zero-shot governance" — domain-agnostic foundation models intervening in decisions, wrapped in thin domain-specific scaffolds.
- K-12 policy: Stanford's evidence reviews of 818 papers inform K-12 AI policy.
- Assessment policy: Assessment reform policies and Authentic Assessment frameworks represent policy-level responses to AI-enabled cheating. The choice of summative assessment format — oral, proctored, closed-book — is itself an assessment-policy decision (see Summative Assessment).
Policy maturity gap
The knowledge base documents that institutional AI policies lag behind actual AI use. Educational Development programs, Teacher AI Competency frameworks, and AI Regulation in Education all require coherent policy foundations. Large-scale field evidence (Strömberg, Lei, & Wu 2026) shows that the learning losses from homework outsourcing go largely unnoticed because individual subject teachers and students rarely connect the decline to AI use — a gap that evidence-informed policy (e.g., weighting closed-book assessment, informing students of long-run costs, monitoring inputs rather than outputs) can address.
Advisory documents outnumber binding policy. A census of every accredited program in one professional field shows that the maturity gap is one of form as well as timing. Eldredge et al. (2026) collected AI policy and guidance documents from all 48 CAHIIM-accredited health informatics and health information management master's programs in the United States and found that 40 (83%) had at least one publicly available AI-related document while 8 had none, but that the documents were mostly guidance rather than enforceable rules: 21 guidelines (53%) against 7 formal policies (18%). Their content centered on academic integrity and acceptable use, with privacy, intellectual property, and regulatory concepts (HIPAA, FERPA, research compliance) appearing far less often, and topic modeling returned the same student-conduct emphasis. Neither document type nor intended audience varied by delivery mode (Fisher's exact P = .85 and P = .71). The authors read this as evidence that academic program policy is a distinct activity from curriculum design and workforce competency development, and argue that accrediting bodies could reduce the resulting variation by providing AI policy frameworks that integrate academic integrity, data ethics, and equitable access.
A sector making policy without an evidence base. The maturity gap also appears as an evidence problem in professional programs. Gutowski and Hurley (2025) scored the institutional policies of ABA-approved US law schools on five dimensions — prohibitiveness, permissiveness, educational integration, transparency and accountability, and depth — and found most schools taking generally prohibitive positions while reserving discretion to individual instructors and committing to review the rules as the technology moves. They attribute the caution to time pressure rather than evidence: the ABA's 2024 survey drew responses from only about 15% of accredited schools, and the authors report no consensus on whether AI use must be disclosed or how it should be cited. Their conclusion is procedural — build policy with faculty rather than for them, train proactively, and treat policy generation as a recurring event rather than a one-time act — which makes Legal Education one of the clearest cases of the policy-versus-governance distinction this page draws.
Instructors regulate; institutions mostly do not. The maturity gap has a second face, visible from the teaching side. Watson and Rainie (2026) surveyed 1,057 US college and university faculty and found that rules have been written far faster below the institution than at it: 87% of respondents created their own assignment-level policies for student AI use, against 35% who said their department has written guidelines and 48% who said their institution has. The structural response behind those documents is thin — 55% reported a task force or oversight group, 37% new AI-focused classes, 17% an AI major or minor, 16% new academic leadership offices, and only 13% adoption of AI literacy as a general education outcome — while 68% said their schools had not prepared faculty to teach with the tools. Because students meet the resulting patchwork as a single institutional regime, the finding is direct evidence for this page's distinction between a policy document and the governance that makes it consistent. (The authors describe the sample as non-scientific and not generalizable, so it is best read as the sector's expressed concerns.)
Task-level regulation is the emerging pattern. A large-scale longitudinal study of 31,000+ course syllabi (2021–2025) at a large public research university (Chirikov 2026) shows how instructors actually regulate AI in practice: explicit AI regulation grew from near zero to 55% of courses by Fall 2025, but the direction shifted from restrictive toward permissive, and instructors increasingly differentiated by task type — restricting AI for drafting/reasoning (displacement-risk tasks) while permitting it for editing/proofreading and study support (augmentation tasks). Framing also shifted from academic integrity (63%→49% of syllabi) toward learning impact (1%→29%). This task-based pattern — built on the labor-economics mechanisms of task displacement, augmentation, and reinstatement — offers a more granular alternative to blanket adoption-or-ban policies and is a direct empirical anchor for the policy-vs-governance distinction above.
Instructor-level policy design as an alignment exercise. Task-level regulation has a design counterpart below the syllabus. McCorkle's (2025) design case derives a course's allowed and unallowed GenAI uses from "what, specifically, am I assessing?" — inventorying every task in a project, mapping each task to a learning objective, pairing each with an emerging GenAI workforce competency (Career Development and Readiness), and then deciding which concern takes priority: the need to assess student performance or the value of building the competency. The resulting policy differs task by task inside a single project — brainstorming topics and curating images permitted, composing learning objectives and designing slides not — with the rationale written directly to students. The case also documents why blanket prohibitions fail in practice: students who do not see themselves as dishonest read a prohibition as not applying to them, and the policy becomes inequitable (Equity) when expectations are left implicit. It complements the top-down policy levels above by showing how a policy's content can be generated from Assessment alignment at the course and assignment level.
The boundary–evidence gap in assessment policy. A 30-university audit of public GenAI assessment guidance (Yao 2026) finds that institutional policies are better at classifying AI use than at explaining what evidence of learning remains valid under each class: the mean delegation-boundary score (2.47/4) exceeded the mean evidence-standard score (1.89/4), safeguards were sparse (2.75 of 8), and guidance was clearest for final-output substitution. The framework of cognitive stewardship argues that policies must make the certification logic visible — what learners may delegate, what they must still demonstrate, and how institutions protect fair evidence — rather than merely monitor AI use.
Over-inclusive prohibition language and misapplied rules. A rule's scope is a policy decision with consequences beyond compliance. Wright (2026) argues that the blanket prohibitions written across higher education since 2023 bar "generative AI" without the technical precision to separate content generation from AI-powered format conversion — optical character recognition, handwritten text recognition and speech-to-text are recognition technologies that infer what a student already wrote rather than producing new content. On this reading, sanctioning transcription-only use is best characterized as policy misapplication rather than misconduct, and the cost of imprecision falls unevenly: equity-exposed and disabled students who rely on those tools carry disproportionate false-positive risk, an exposure that legal issues and risks treats as a reasonable-adjustment question rather than an integrity one.
A purpose-based boundary instead of a tool list. Li (2026) supplies the policy content that over-inclusive prohibition language lacks: a support-versus-substitution line defined by function and the assessment construct rather than by naming tools, on the argument that tool lists go obsolete and encourage compliance by interface. Permitted support is a surface intervention that adds no ideas, no sources and no material re-ordering of analysis; substitution is generating arguments or counterarguments, applying disciplinary rules to facts, restructuring an analytical sequence, or generating citations for inclusion. The framework is delivered as three instruments — working definitions with shared quick tests so two markers characterize the same conduct alike, calibrated disclosure templates ranging from nothing for embedded low-risk functions to a one-line statement for external language editing and roughly 80–120 words where limited ideation is expressly authorized, and a decision rubric that names weak proxies (polished language or sudden fluency, non-native phrasing, single detection scores) beside the evidence types that can carry weight. It also states the conditions under which scaled permissive regimes such as the AI Assessment Scale remain workable: use expressly authorized for the specific task, the construct re-specified so students and markers know what is assessed, and disclosure kept feasible.
Staged automation with mandatory human oversight. Uruguay's Acredita EB national lower-secondary accreditation test offers a concrete model of oversight-preserving automation in assessment policy: multiple-choice sections and existing procedures stay unchanged, AI is used only for the writing section inside a mandatory calibration stage on 50–100 texts, and two pruning rules decide which results require expert review — candidates who cannot pass on the other two sections need no writing review, and candidates Proficient in all three sections pass with only 0.2%/0.6% observed residual risk. The authors estimate the design could cut written responses needing full human scoring by at least 50% while keeping the guiding principle that no final decision about test outcomes is made without appropriate human oversight. This complements the boundary–evidence work above by showing a policy that specifies not only where AI is permitted but which human evidence and review remain mandatory.
The adversarial side of automated grading. A policy that routes grading through an AI tool inherits an attack surface. Humble (2026) red-teamed an institutional grading workflow by hiding five indirect prompt injections inside the files of a synthetic submission that Microsoft Copilot (GPT-5.2) had graded as fail on all six baseline runs. Two strategies raised the grade with no visible warning to the user — reported attack success rates of 100% (9 of 9 iterations) and 94% (17 of 18) — and the tool both silently disabled a chat after blocking the simplest attack and, in one iteration, announced it would ignore embedded instructions before raising the grade on the next six runs. A grade obtained through a hidden instruction carries no validity claim, and the same technique could degrade a submission with no durable trace, so the paper's asks are policy-level rather than technical: sector-level guidance, professional development, restricted AI use with human review for high-stakes work, and standardized, domain-agnostic testing of prompt-injection resilience so the attack surface is measured rather than assumed. It is the adversarial counterpart to the oversight designs above: where those specify which human review is mandatory, this specifies a class of manipulation that review exists to catch but has no reliable signal for.
A value/norm matrix as a policy foundation. Agarwal et al. (2026), a systematic review of 25 articles, consolidate AIED ethics into six main ethical values (non-discrimination, data stewardship, human oversight, goodwill, explicability, educational aptness) and map the ethical norms extracted from the literature onto a stakeholder-by-value matrix. The review positions the matrix as a foundation for building detailed ethical frameworks and regulation for AIED, giving educational institutions, developers, and regulators concrete norms to implement specific values. It finds goodwill norms aimed at regulators are far more numerous (nine) than for any other stakeholder set, signaling regulators' role in ensuring AIED benefits learners through policy and legislation — a concrete, value-anchored starting point for the policy-vs-governance machinery this page describes.
Claims that travel further than their qualifications. A pilot announcement is a policy instrument too, and its reporting standards are part of the policy. Restrepo Morales et al. (2026) analyze the September 2026 El Salvador episode in which an AI-tutoring pilot in 171 public schools was reported as reaching results comparable to Germany and Sweden: against the German PISA 2022 mathematics average, the top 5.5% of the national achievement distribution would have matched the benchmark with no learning gain at all, and the pilot's assessment covered 7.0 students per school against 25.4 in the national survey the same year, so the published evidence cannot exclude selection as the explanation. Their proposal is a six-item reporting standard — the participating schools' baseline, the sampling protocol (eligible students, assessed students, identification rule, participation rate), disaggregated scores with standard errors, the operational date of each reform component, overlap with the representative national sample, and the qualification stated in the same document, post or paragraph as the claim — motivated by the disclosure arithmetic of the announcement itself: the World Bank post carrying the claim recorded about 441,000 views against about 11,000 for the post carrying the qualification four places later in the same thread. The policy lesson is that where a system uses school-level results as evidence about itself, the format of disclosure decides which statement becomes the public fact.
Policy vs. governance
Policy and governance are closely related but distinct, and keeping them apart matters for understanding the knowledge base's research.
- Policy is the content — what is decided. A policy is a formal rule, principle, or statement: what AI use is allowed, prohibited, or required; what must be disclosed; what assessment formats are permitted. It exists as a documented artifact (legislation, institutional guideline, syllabus statement) and answers "what are the rules?"
- Governance is the machinery — how it is decided, implemented, and enforced. Governance encompasses the institutional structures, norms, and accountability mechanisms that produce, carry out, and monitor policy: who sets the rules, how they are communicated and resourced, how compliance is enforced and challenged, and how they are revised as AI evolves. It answers "who decides, and how do the rules actually take effect?"
- They are interdependent. Policy without governance is unenforced — a written rule no one owns, monitors, or updates. Governance without policy lacks direction — structures that administer nothing in particular. The knowledge base's research repeatedly shows that the two must be built together: a policy that only classifies AI use without governance to specify evidence, safeguards, and revision processes remains weak in practice (the cognitive-stewardship audit), and governance that merely monitors without clear policy risks surveillance without fairness (see AI governance).
The practical test that separates them: a policy can be read on paper, but governance is observed in whether the rule is implemented, enforced, and adapted. This is why AI Governance extends AI Regulation in Education and policy into institutions, and why the knowledge base treats assessment-format choices (Summative Assessment) as policy decisions that only become effective through governance structures like review boards, declaration frameworks, and appeal routes.
Classroom policy as the interpretive frontier of governance. The machinery of governance does not stop at the institutional document; the teachers who read and adapt it are the last link in the chain, and Nash & Burriss (2026) show what that link looks like in practice. Their preservice teachers expected to work in districts with AI policies yet still had to author classroom-specific rules for their own students — an essential adaptation skill — and the authors argue that teachers can decline to adopt AI for specific reading and writing tasks without ignoring it, with principled refusal distinguishable from uncritical rejection. Keeping that distinction viable is itself a governance task: districts and preparation programs need the guidance, professional development, and policy infrastructure that let teachers refuse particular uses without being framed as behind the times.
Governance indicators for the authentication problem. Coates, Croucher and Calderon (2025) supply the measurement layer this policy-versus-governance distinction implies, treating Academic Integrity in the GenAI era as a governance problem before a detection problem and judging contemporary academic governance resilient but "not well positioned or poised" to protect the authentication of student Assessment. Their framework holds 130 governance questions under eight dimensions — Designing, Developing, Training, Implementing, Analyzing, Reporting, Evaluating and Improving — and the items are questions about institutional machinery rather than psychometric scales: whether the institution's top-most board or council receives updates on assessment processes and outcomes, whether key performance indicators cover assessment quality, what percentage of students are known individually by the teachers who assess them, whether extreme low or high marks are cross-checked, and whether there is a simple route for referring contract cheating cases. The accompanying reform program targets governance architectures, the people who hold governance roles, and the technologies and resources supporting assessment, and the authors argue that little institutional development pays out without external affordance from regulation, benchmarking and cross-institutional competition — the same regulatory pressure the AI Regulation in Education page treats as the binding constraint on governance reform.
The policy deficit in AI × SEL research
A systematic review of 65 papers at the intersection of AI and social-emotional learning (Tran, Liu & Nguyen 2026) documents a "policy deficit": nearly three-quarters of studies state no policy implications, and those that do often lack actor-oriented specificity. The review finds policy engagement correlates with publication venue, reflecting academic incentives that reward technical novelty over AI Governance and AI Regulation in Education. It proposes a "WH-question" framework (Who, What, Why, When/Where, How) and a shift from "implication-as-afterthought" to "implication-as-methodology" — treating policy articulation as a design constraint of research rather than a post-hoc add-on.
Connected Concepts
- Pedagogical Partnerships — Pedagogical Partnerships
- AI Regulation in Education
- AI Governance
- Educational Development
- Equity
- Higher Education
- K-12
- Academic Integrity
- Ethics
- Teacher AI Competency
- Framing AI Use for Students
- Summative Assessment
- Chemistry Education — Chemistry education and AI: labs, formative assessment, LLM limits, philosophy of experimentation
- Biology Education — Biology education and AI: lab teaching assistants, AI literacy in biology, critical thinking, specialized tools
- Stakeholders — Umbrella: people and audiences in AI education (learners, teachers, designers, administrators, policymakers)
Connected Articles
- Writing the Rules for Generative Machines: Tensions and Entanglements in Preservice Teachers' Classroom AI Policies — Preservice English teachers' classroom AI policies: what they permitted, limited, and banned (Nash & Burriss 2026)
- A Human-in-the-Loop Framework for AI-Assisted Scoring in Large-Scale Writing Assessment — A Human-in-the-Loop Framework for AI-Assisted Scoring in Large-Scale Writing Assessment
- Designing an Aligned Generative AI Course Policy: An Equitable and Transparent Learner-Centered Approach — Deriving allowed and unallowed GenAI uses task by task from what is assessed (McCorkle 2025)
- How Instructors Regulate AI in College: Evidence from 31,000 Course Syllabi — How instructors regulate AI across 31,000 course syllabi (Chirikov 2026)
- Artificial Intelligence and Grade Inflation — AI task displacement as a mechanism of grade inflation (Chirikov 2026)
- The AI Adaptation Gap in Higher Education: Students, Faculty, and Administrative Staff — The AI Adaptation Gap in Higher Education
- Governing generative AI in higher education: a global Delphi study on policy and practice — Global Delphi on governing generative AI in higher education
- Governing generative AI in higher education: Emerging policy approaches and support ecosystems at innovative U.S. universities — Guidance over policy: instructor-set syllabus rules and a four-unit support ecosystem across 50 innovative US universities (Qian 2026)
- Forging ahead or proceeding with caution: Developing policy for generative artificial intelligence in legal education — Five-factor comparison of law school GenAI policies in a sector making policy without an evidence base (Gutowski & Hurley 2025)
- Transcription is not generation: Distinguishing non-generative AI tool use from academic misconduct in higher education assessment — Transcription is not generation: over-inclusive AI rules, format conversion and disability accommodation (Wright 2026)
- Anticipatory governance and leadership for AI implementation in higher education: A scoping review — Anticipatory governance for AI in higher education (scoping review)
- Institutional approaches to artificial intelligence policy and guidance in health informatics and information management education: emerging trends and inconsistencies — AI policy documents across all 48 accredited health informatics master's programs: mostly guidance, centered on academic integrity (Eldredge et al. 2026)
- What Does the Credential Still Certify? Cognitive Stewardship for AI-Mediated Education — Cognitive stewardship for AI-mediated assessment (30-university policy audit)
- Generative Artificial Intelligence Policy: A Qualitative UNESCO Framework Analysis
- A Comparative Analysis of Institutional and Course Generative AI Policies within Higher Education: Implications for Instruction in Computing Education
- A Framework for Institutional Change in the Age of AI
- A bit of chaos and madness: The AI Assessment Scale and the work of assessment reform
- The Evidence Base on AI in K-12: A 2026 Review
- Artificial Intelligence in UK Higher Educational Policy and Institutional Decision Making
- Reassessing Academic Integrity in the Age of AI: A Systematic Literature Review on AI and Academic Integrity — Call for explicit, co-developed AI-use policies
- Young People, Learning, and Generative AI: A Rapid Literature Review and Implications for PreK-12 Education — Move beyond adoption-or-ban; staged, developmentally responsive guidance
- The Policy Deficit in AI × Social-Emotional Learning Research — The Policy Deficit in AI × SEL Research
- Identifying the ethical values and norms for artificial intelligence in education: A systematic literature review — Ethical values and norms for AI in education
- Zero-Shot Governance: General-Purpose AI in Policy — Zero-shot governance: general-purpose AI in policy (Perrotta 2026)
- How much selection would be enough? Bounding the learning claim of El Salvador's artificial intelligence tutoring pilot — Bounding the learning claim of El Salvador's AI tutoring pilot (Restrepo Morales et al. 2026)
- The AI Challenge: How college faculty assess the present and future of higher education in the age of AI — 1,057 US faculty: 87% write their own assignment rules against thin institutional and departmental guidelines (Watson & Rainie 2026)
- GenAI assessment and language equity: Drawing the line between support and substitution — A purpose-based support–substitution boundary with calibrated disclosure and decision rubrics (Li 2026)
- Ethical implications of prompt injection in AI-mediated grading: An adversarial red-team evaluation — Prompt injection in AI-mediated grading: grades changed undetected, and the policy-level response (Humble 2026)
- Governing academic integrity: Ensuring the authenticity of higher thinking in the era of generative artificial intelligence — 130 governance indicators for authenticating assessment, and the external pressure reform needs (Coates, Croucher & Calderon 2025)
- A Workshop Series for Effective Use of AI in Uncertain Times: Building a Physics Faculty Learning Community — A Workshop Series for Effective Use of AI in Uncertain Times: Building a Physics Faculty Learning Community
- Who Acts, Who Knows, Who Answers? A Corpus-Assisted Discourse Analysis of Agency, Epistemic Responsibility, and Accountability in Generative AI Higher Education Research — Who Acts, Who Knows, Who Answers? A Corpus-Assisted Discourse Analysis of Agency, Epistemic Responsibility, and Accountability in Generative AI Higher Education Research
- "We'll Fix It Later": Education, AI, and the Deferral of Student Privacy in EdTech — "We'll Fix It Later": Education, AI, and the Deferral of Student Privacy in EdTech