AI in K-12 Evidence Base

Created: 2026-05-07 | Tags: k-12rctefficacy-studylearning-gainsllmadaptive-learning
๐Ÿ“„ Full text: Stanford SCALE ยท local
As of October 2025, only 20 of 818 papers in the AI Hub Research Repository meet standards for strong causal inference (RCTs or QEDs) on AI in education โ€” and zero examine U.S. K-12 student settings.^stanford-evidence-base-ai-k12-2026

The Evidence Gap

The field is growing explosively (from 28 relevant papers in Jan 2023 to 818 by Oct 2025) but remains thin on causal evidence:

The review authors (Stanford SCALE, 2026) applied What Works Clearinghouse (2025) standards for quality review, using a two-step LLM pre-screening followed by human review.

What the Causal Literature Says

Out of 20 high-quality causal papers, outcomes break down as:

Education levels in causal papers:

Learning Science Lens

The review frames findings through six established principles. See ai-learning-transfer for the critical tension between immediate performance and durable learning.

Principle Implication for AI Tools
Cognitive load AI reduces extraneous load but may also suppress germane (productive) load.
Zone of proximal development General-purpose AI often operates outside the ZPD by completing work for students. Tutoring-specific scaffolds target readiness more precisely.
Transfer Unclear whether AI-assisted practice produces durable knowledge or merely tool-dependent performance.
Metacognition AI completing tasks reduces opportunities for students to monitor their own understanding.
Expertise reversal Novices need guidance; experts need independence. Effective AI must adapt to learner expertise.
Desirable difficulties Easier practice with AI feels better but may weaken long-term retention.

Key Open Questions

1. Do AI-assisted gains persist when the tool is removed? 2. Does general-purpose chatbot use widen achievement gaps? 3. What pedagogical guardrails are necessary to preserve metacognition? See the new evidence from Scheu et al. (2026) that structured courses increase metacognitive competence while raw LLM assistants do not. 4. When will the first U.S. K-12 RCTs on LLM tools emerge?

Related Pages

-