---
source_url: https://scale.stanford.edu/sites/default/files/The%20Evidence%20Base%20on%20AI%20in%20K-12%20Report.pdf
ingested: 2026-05-07
sha256: 34464b4a511b295621c7f2478e3732159b9433bcaa3921cf82ea28dc99549356
---

# The Evidence Base on AI in K-12: A 2026 Review — Summary

**Source:** Stanford SCALE Initiative, AI Hub for Education
**Scope:** Analysis of 818 papers (as of Oct 2025); only 20 provide strong causal evidence (RCTs/QEDs).

---

## Critical Context: The Evidence Gap

> "Research on how AI impacts K-12 students and educators is still extremely limited."

- **818 papers** in the AI Hub Research Repository (Oct 2025); only **20** met standards for strong causal inference.
- **Zero** high-quality causal studies examine U.S. K-12 *student* settings; very few exist for U.S. K-12 *educators*.
- Most causal research is international, conducted in postsecondary settings, short-term (e.g., one 20-minute session), and focused on immediate outcomes.

---

## Methodology at a Glance

- **Repository sources:** 87% from arXiv preprints; updates monthly.
- **AI definition:** Focuses on LLM-era AI/ML tools; excludes pre-LLM rule-based intelligent tutoring systems.
- **Quality review:** Two-step process—LLM pre-screening followed by human review using *What Works Clearinghouse (2025)* standards.
- **Limitations:** Repository is preprint-heavy; search terms limited to "education" + "AI" or "artificial intelligence."

---

## State of the Field: Characteristics of Research

| Characteristic | Full Repository (818) | Causal Impact Papers (20) |
|---|---|---|
| **Primary users studied** | 59% students; 48% educators | 70% students; 40% educators |
| **Top outcomes** | 50% other academic; 17% math | 35% math; 25% other academic; 20% literacy; 15% social-emotional |
| **Education level** | 64% postsecondary | ~5% postsecondary; 45% high school |
| **Study design** | 46% descriptive; 46% technical/computational; 8% RCT; 5% QED | 90% RCT; ~20% QED |

*Growth note:* Only 28 relevant papers existed in Jan 2023; the repository doubled between Jan–Sept 2025.

---

## Learning Science Framework

Findings are interpreted through established principles:

| Principle | AI Opportunity & Risk |
|---|---|
| **Cognitive Load** | AI can reduce extraneous load but also *germane* (productive) load essential for learning. |
| **Zone of Proximal Development** | General-purpose AI may operate outside this zone by doing the work for students. Tutoring-specific hints/scaffolds better target learner readiness. |
| **Transfer of Learning** | Unclear if AI-assisted practice yields durable knowledge or merely tool-dependent performance. |
| **Metacognition** | AI completing tasks reduces opportunities for students to monitor understanding and select strategies. |
| **Expertise Reversal** | Novices benefit from guidance; experts benefit from independence. Effective AI adapts to expertise. |
| **Desirable Difficulties** | Easier practice with AI feels better but may harm long-term retention and transfer. |

---

## Key Findings: Students
*Based on 14 causal studies (7 university; 7 high school international). **None in U.S. K-12.***

### 1. Immediate Gains with Access
AI significantly improves performance **while students use it**:
- Math proofs, programming, economics exams, physics, and argumentative writing all show gains during AI-supported practice.

### 2. Short-Term Boost, Uncertain Transfer
Effects are **mixed or negative** when AI is removed:

> "AI tools may help students complete tasks more successfully in the moment, but those gains do not always persist when students are later asked to perform independently."

- **Bastani et al. (2025):** High schoolers using a general-purpose chatbot for math practice performed **~17% worse** on closed-book final exams than peers with no AI, despite higher practice grades.
- **Chen et al. (2025):** LLM-Tutor improved homework but **did not improve unassisted exam scores**.
- **Lehmann et al. (2025):** General-purpose AI for programming increased topics covered but **harmed understanding** and widened achievement gaps for low-prior-knowledge students.
- **Stadler et al. (2024):** General-purpose AI reduced cognitive load but produced **lower-quality reasoning and argumentation** vs. traditional search.
- **Kosmyna et al. (2025):** AI essay assistance led to **83% of participants failing to recall a quote** from their own essay, vs. 11% for non-AI users.

### 3. Easier Doesn't Mean Better
> "While AI can reduce perceived difficulty and increase fluency, this may come at the cost of reduced independent reasoning and weaker knowledge acquisition."

- Students report greater enjoyment and reduced cognitive burden (Becker et al., 2025; Stadler et al., 2024).
- However, reduced effort can undermine deeper learning. **Kreijkes et al. (2026)** found retention improved only when AI use was paired with traditional strategies like note-taking.

### 4. Pedagogical Design Matters
Tutoring-specific tools consistently outperform general-purpose chatbots:

- **Bastani et al.:** A tutoring-specific chatbot with pedagogical guardrails (hints, step-by-step reasoning) mitigated the exam drop; general-purpose GPT Base caused it.

---

> **Note:** This extraction was truncated by the web extraction service. Additional sections (educator findings, policy implications, recommendations) may exist in the original PDF. Re-ingest with full PDF parsing for complete coverage.
