On this page

Synthesis: This scoping review systematically maps 474 studies (January 2020 – May 2026) of generative-AI-powered agentic AI in education, the most comprehensive synthesis of the field to date. Analyzing publication characteristics, study designs, agent roles, model architectures, six dimensions of agentic capability, and the extent of educational theory integration, it finds rapid post-2025 expansion concentrated in higher education, STEM, and text-based tutoring, but modest capability levels: systems rarely exhibit strong tool orchestration, embedded AI Governance, or persistent memory. The review exposes a disciplinary divide — only 138 of 474 studies (29%) drew on educational theory — and converges on priorities of longitudinal validation, stronger pedagogical grounding, and more governed adoption of frontier agent infrastructures.

Summary

This scoping review systematically maps 474 studies (January 2020 – May 2026) on generative AI-powered agentic AI in education, providing the most comprehensive synthesis of the field to date. The authors analyze publication characteristics, study designs, agent roles, AI models/architectures, six dimensions of agentic capability, and the extent of educational theory integration.

Key Findings

1. Rapid Expansion Since 2025

The field has grown explosively, but the literature is dominated by conference papers concentrated in higher education, STEM disciplines, and text-based tutoring scenarios. This mirrors the general trajectory of AI in Education research, but with a specific agentic inflection point in 2025.

2. Technology Stack: GPT + LangChain Dominate

GPT-series models and LangChain are the most widely adopted technologies. Notably, OpenClaw and other frontier agent paradigms (governed tool orchestration, persistent memory, long-horizon planning, multi-agent coordination) remain largely absent from educational research — revealing a significant technology–application gap. This stands in contrast to the vision articulated in Agentic AI.

3. Agentic Capabilities Remain Modest

Across the six capability dimensions analyzed:

  • Single-task autonomy — commonly demonstrated ✓
  • Sequential planning — increasingly present ✓
  • Multi-agent collaboration — growing ✓
  • Strong tool orchestration — rarely exhibited ✗
  • Robust embedded governance — rarely exhibited ✗
  • Persistent memory / long-horizon planning — largely absent ✗

This maps closely to the four-paradigm framework in Evolution of AI in Education: Agentic Workflows (reflection, planning, tool use, multi-agent collaboration), where the reviewed systems tend to cluster in the first two paradigms while falling short on the more advanced ones.

4. Theoretical Grounding is Limited

Only 138 of 474 studies (29%) explicitly drew on educational theory, revealing a clear disciplinary divide between technically oriented research (CS/engineering) and pedagogically oriented work (education/learning sciences). This echoes broader concerns about the gap between technological capability and pedagogical intentionality.

5. Methodological Limitations

Most studies rely on small-scale, short-term designs. Longitudinal and real-world validation studies are rare, limiting the evidence base for claims about effectiveness. The review calls for more rigorous efficacy-study designs and attention to Student Experience beyond immediate performance metrics.

6. Six Dimensions of Agentic Capability (the Review's Analytical Framework)

Dimension Description Status in Reviewed Studies
Task Autonomy Independent task initiation, planning, completion Common
Goal-Directed Reasoning Strategy selection and adaptation to context Emerging
Memory & Context Awareness Using interaction history and learner profiles Limited
Planning & Sequencing Multi-step plan formulation and execution Growing
Tool Orchestration Invoking and coordinating external tools/resources Rare
Governance & Oversight Auditable action, safety constraints, human-in-the-loop Rare

The governance gap is particularly concerning given frameworks like Human-in-the-Loop, which emphasize that educational AI systems require robust oversight mechanisms — not just technical capability.

Priority Research Directions

The review identifies several converging priorities:

  1. Longitudinal and real-world validation — moving beyond short-term lab studies
  2. Stronger pedagogical grounding — bridging the CS/education disciplinary divide
  3. Governed adoption of emerging agent infrastructures — particularly tool orchestration and multi-agent coordination
  4. Systematic integration of Ethics and human oversight — connecting to Equity and Academic Integrity concerns
  5. Expanding beyond STEM and higher education — into K-12, language learning, special education, and Workplace Learning contexts

OpenClaw as an Analytical Lens

The review uses OpenClaw (Steinberger, 2026) — the fastest-growing Open Source AI project in early 2026 — as an illustrative reference point for the "frontier agent paradigm": systems that feature governed tool orchestration via MCP, persistent memory, long-horizon planning, multi-agent coordination, and auditable action. The finding that these capabilities are largely absent from educational agentic systems is the review's most striking technology–application gap. While the authors are careful not to position OpenClaw as a normative target, its feature set serves as a useful Benchmark for assessing how far educational systems lag behind general-purpose agentic infrastructure.

What this means for practice

  • Designers. Test whether governed tool orchestration, persistent memory, and long-horizon planning can serve learning goals rather than assuming simple chatbot interaction is the ceiling — the review's technology–application gap shows the field uses only a narrow slice of available agentic capability.
  • Researchers. Pair technical system-building with explicit learning theory and instructional design so automation supports rather than replaces learner cognition; only 138 of the 474 reviewed studies (29%) drew on educational theory.
  • Researchers. Treat effectiveness claims cautiously and design longitudinal, real-world validation studies, since most reviewed work uses small-scale, short-term, single-context samples.
  • Administrators. Make AI Governance, human oversight, equity, and Academic Integrity core design constraints rather than afterthoughts as systems scale toward greater autonomy and multi-agent coordination.
  • Instructors. Scrutinize any agentic tutor before adoption for auditable action and human-in-the-loop checkpoints, because robust embedded governance was rare across the 474 systems reviewed.

Limitations

  • As a scoping review, the study conducted no meta-analysis and did not assess risk of bias in individual studies, so it maps the field's shape but cannot establish effectiveness.
  • The search was restricted to titles, abstracts, and keywords, and screening used an LLM-assisted strategy with two human reviewers, leaving relevant studies retrievable only through full-text indexing outside the corpus.
  • The evidence base itself is dominated by small-scale, short-term designs and post-2025 publication (278 of 474 studies in 2025, 146 in January–May 2026), limiting the maturity of any cumulative claim.
  • First authors are heavily concentrated in China (150, 31.6%) and the United States (93, 19.6%) — 51.2% combined — which the authors flag as a threat to generalizability across educational systems.

Citation

Wang, N., Zou, D., Xie, H., & Qin, S. J. (2026). A scoping review of generative AI-powered agentic AI in education: Research landscape, agentic capabilities, and insights from the frontier agent paradigm, exemplified by OpenClaw.

Embed this page

Copy the code below to embed a chromeless version of this page in a learning management system or other website. The embedded view hides the site header, navigation, and footer.