On this page

Synthesis: Raza & Farooq (2025) content-analyze 100+ peer-reviewed AI-in-education studies from 2020–2025, the era of mainstream LLMs and generative AI. They organize the literature into three interrelated layers — the genome layer (evaluation practices, leadership capacity, low-friction tools, shared norms, algorithmic schemas), the cognitive layer (predictive analytics, personalized learning via a measure–model–adapt loop, Multimodal AI sensing, discourse and affect analysis), and the symbiotic layer (end-to-end learning platforms, smart classrooms, process automation, GenAI copilots). Across studies they summarize effects on learning, engagement, teacher workload, and adoption, and distill three forward trends: (1) human–AI co-orchestration as the default classroom pattern; (2) privacy-preserving, edge/federated AI for sensitive student data; and (3) authentic, continuous assessment via Multimodal AI analytics and generative simulations. The core message is practical: invest first in people and workflows ("plumbing"), then in models that earn trust, and only then in platforms that scale what already works.

Three-Layer Framework

This review reads the last five years of AI-in-education through a three-layer lens that moves from preconditions, to mechanisms, to systems:

This structure lets the authors evaluate studies on more than effect size: whether results are actionable (who does what, when), credible (transparent and fair), and durable (likely to hold under real constraints).

Key Findings

  1. 100+ peer-reviewed papers (2020–2025) were coded with a reproducible technique/contribution/outcome scheme; the corpus skews toward higher education and the Global North, limiting generalizability.
  2. Genome layer: successful adoption depends on formative, actionable Feedback for teachers, leaders who act on evidence, tools that fit existing routines, and shared norms (e.g. Dignum's ART framework) — not on chasing advanced models.
  3. Cognitive layer: learning gains are most reliable when instruction follows a measure–model–adapt loop — make learning visible, model risk or progress with interpretable models, and close the loop quickly; edge/fog and on-device computation reduce latency and keep adaptations meaningful.
  4. Symbiotic layer: gains persist where sensing, analytics, and teacher action are tightly coupled; GenAI works best as a "strong junior assistant" — strong at drafts and idea generation, inconsistent at judgment — so humans must stay in the loop.
  5. Three forward trends: (1) human–AI co-orchestration with explicit handoff protocols; (2) privacy-preserving, edge/federated AI for sensitive student data; (3) authentic, continuous assessment via Multimodal AI analytics and generative simulations.

Genome Layer: Foundations

At the genome tier, the point isn't to chase advanced models; it's to get the groundwork right. Assessment only helps when it is folded back into teaching as a routine, not an audit — Huang et al. show evaluation works when instructors get usable feedback on performance, attitudes, and satisfaction rather than punitive scores. Xu extends this to the institutional layer: optimization can route resources and attention, but only if the organization is set up to act. Parapadakis offers a caution that automated survey analytics are fast yet fragile — without Guardrails, speed can harden weak inferences. That is why leadership readiness matters: Anysiadou and Gkliati find only moderate preparedness among school leaders, while Lu shows how long-term university–industry partnerships build the muscle memory to adopt AI responsibly. The foundation isn't dashboards; it's decision pathways.

Early algorithmic work is best read as proof-of-possibility that supports, rather than drives, Pedagogies and Teaching Strategies. Kim and Kim and Jiang et al. show targeted gains in language tutoring and oral-English correction, while Zafari et al. note that across K-12 STEM, machine learning (ML) and Intelligent Tutoring dominate because they are adaptable and close to teacher needs. Leitner et al.'s ARIN-561 shows students need narrative entry points to AI concepts. Williamson names the drift toward data-intensive "AI-based learning science", which raises social signals: Wen et al. capture public ambivalence about classroom ChatGPT, and Xie and Wang find no IQ/memory harm — useful, but not a license for overreach. Treat algorithms as Scaffolding: their value rises when they disappear into better routines, not when they become the center of the lesson.

Methodological hygiene and AI awareness complete the layer. Doroudi reminds us AI and the learning sciences have co-evolved; Yang et al. map the field's themes via bibliometrics and knowledge graphs; Saltman warns that without reflection AI can fuel teacher de-skilling and neoliberal pressures. Studies across learners show a wide spread — low awareness among science students, positive attitudes in Indian secondary schools, optimism with shallow understanding among younger learners. Educators report openness but lack training, which is why Dignum's ART framework (Accountability, Responsibility, Transparency) is a useful anchor. Culture beats tooling: shared language, clear principles, and basic literacy make the higher layers possible.

Cognitive Layer: Mechanisms

At the cognitive tier, AI stops being background plumbing and starts shaping day-to-day instructional decisions. A useful reading is a measure–model–adapt loop: schools make learning visible, model risk or progress, then adapt instruction or supports. Prediction can turn Assessment from a few high-stakes snapshots into ongoing support — Yuan's CTQAS gives teachers real-time monitoring, Pallathadka et al. predict grades (SVM reaching 88% accuracy), and early-warning work (Villegas-Ch et al.; da Conceição et al. with interpretable ML) surfaces practical dropout predictors like GPA, age, and attendance. Interpretable models are not "nice to have"; they are what let leaders justify and target support.

Models embedded in teaching workflows tailor content, pace, and feedback. Sun et al.'s DL-OIET improves online English via personalized recommendations; Ma et al. pair hybrid optimization with student clustering; Li proposes a multi-agent architecture; Han et al. use fog computing and hierarchical Q-learning for lower latency. In physical education, Li and Wang use wearable sensors for real-time feedback. The wins come from short feedback cycles and proximity: sensing, deciding, and responding near the learner reduces lag and keeps adaptations meaningful.

Without good signals, personalization guesses. Chu et al. extract discourse indicators from video and find content similarity predicts dissatisfaction; Liu and Zou model live classroom interactions; Zhao and Yu's ANN supports classroom emotion recognition; Kim et al. show AI can especially boost Creativity for lower-skill students in collaborative art. Intention matters too — using the Theory of Planned Behavior, Chai et al. find Self-Efficacy is a strong predictor of intention to learn AI. Better sensing is not surveillance; it is context that lets you adapt earlier and more precisely, especially for the learners who benefit most.

Reviews and frameworks consolidate what works. Zahariev et al. chart adaptive assessment and Intelligent Tutoring; Holmes et al. argue for human-centric AIED; Cheng et al. emphasize Trust as a precondition. Domain critiques matter: Liu and Afzaal caution that machine translation assists best with human oversight, Gray highlights fairness and surveillance risks, and Filgueiras analyzes platformization and the Beijing Consensus. On the ground, adoption follows predictable routes — institutional AI capability correlates with better performance (Wang et al.), and Innovation Diffusion Theory explains uptake (Almaiah et al.). Classroom-facing work shows where GenAI helps and stalls: Pretorius uses prompt-design to turn GenAI into a mirror for reflective practice, while Li et al. find ChatGPT grading struggles with higher-order thinking relative to teachers.

Symbiotic Layer: Systems

Here the conversation shifts from "smart components" to living systems: platforms, classrooms, policies, and daily academic work where humans and AI co-operate. The deeper message is about infrastructure: treat AI as infrastructure — invest in data flows, teacher workflows, and oversight from day one — rather than a plug-in that chases pilots. Historical and sociotechnical readings remind us scale is path-dependent; values, incentives, and power arrangements shape outcomes.

Real deployments report gains across multi-course platforms, smart-classroom models that marry e-schoolbags with cloud analytics, virtual environments for teacher training, and targeted chatbots for clinical skills. A large quasi-experiment in Russian schools shows performance and engagement lifts while flagging ethics and infrastructure issues. The gains aren't "AI magic"; they are the outcome of good plumbing plus good Pedagogies and Teaching Strategies. Success shows up where sensing, analytics, and teacher action are tightly coupled — if any leg is weak, results fade.

Systems stick because people find them useful and supported. Students adopt AI teaching assistants when they perceive clear utility; AI literacy predicts intention to learn AI; peer networks and leadership make or break institutional adoption. Budget for the "last mile": teacher time, peer champions, and role clarity — adoption follows perceived usefulness plus social proof plus support, not model accuracy alone.

On the ground, GenAI is changing work patterns. For coding, students get faster feedback and broader entry ramps than with legacy forums; faculty use GenAI to draft outcomes and assessments; course-level integrations report achievement and motivation gains when teachers curate and constrain the tools. In distance education, engagement is a force multiplier — AI helps most when learners are already leaning in. Credibility is an outcome of design, not a press release: labeling evidence as "AI-framed" can lower perceived credibility, and in assessment, efficiency claims outpace fairness guarantees. Equity-centered guidance focuses on who benefits, who is burdened, and how to keep classrooms human-safe; economics warn of polarization unless the upside is shared. National readiness varies — some systems lack the digital footing to adopt safely at scale.

Looking across recent scholarship and policy roadmaps, three trajectories appear credible and consequential for the next decade, each reflecting a shift from "AI that does things to learners" toward "AI that works with learners and educators":

1. Human–AI co-orchestration as the default classroom pattern. Teachers and students are less likely to be replaced by AI than partnered with it. Design research on "teacher copilots" is coalescing into staged-rollout frameworks, and field deployments in low-resource contexts show measurable reductions in planning time and stress. The frontier is explicit division of labor across planning, delivery, and assessment — formalizing handoff protocols (what the copilot proposes, what the teacher accepts or edits, and how those edits update future AI behavior) and studying their effects on equity and learning.

2. Privacy-preserving, on-device AI for sensitive learning data. Compliance pressures are pushing computation closer to the learner via federated learning and compact on-device language models that run locally. The next wave will instrument edge-first architectures — Multimodal AI sensing and feedback that never leave the classroom network, with only differentially private summaries shared upstream. Open questions include how much learning signal is lost or gained on device, and how schools govern model updates when each device adapts locally.

3. Authentic assessment via multimodal analytics and generative simulations. Assessment is migrating from end-of-unit tests to in-the-flow evidence captured during authentic tasks. Multimodal learning analytics (MMLA) are maturing toward classroom-ready toolkits, while GenAI powers role-play and case simulations that elicit higher-order skills. The center of gravity shifts from "grading products" to "modeling processes" — assessment for learning, not just of learning.

What this means for practice

  • Researchers. Report total cost of ownership — teacher time saved, implementation cost, compute and energy use — alongside learning and equity outcomes, using the cost–time–quality template the authors propose, because many studies in this corpus report learning effects but omit those figures.
  • Sequence adoption by layer: settle formative assessment and Feedback routines, leadership capacity, low-friction tools, and shared norms before adding predictive or generative components — the review treats those preconditions as what carries the layers above them.
  • Design instruction around a measure–model–adapt loop, and prefer interpretable, edge/fog or on-device models that shorten the feedback cycle and keep a human in the loop where errors are costly.
  • Constrain generative AI to drafts and idea generation under rubrics, exemplars, and human review; the review characterizes it as strong at drafts and inconsistent at judgment.
  • Tie evaluation to decision pathways — who acts, on what evidence, and when — rather than to model accuracy, and set equity targets before deployment so co-orchestration does not simply scale whatever the system already does.

Limitations

  • The corpus covers 100+ peer-reviewed papers published between January 2020 and April 30, 2025 and prioritizes English-language, peer-reviewed education sources, so technical venues, non-English work, and gray literature are underrepresented.
  • Coverage skews toward higher education and the Global North, which the authors state limits generalizability to K–12, vocational, and low-resource contexts.
  • The synthesis codes each study's technique, contribution, and outcome and reports frequency summaries; it pools no effect sizes and reports no formal quality appraisal, so it maps the field rather than estimating effect magnitudes.
  • Many included studies report learning effects but no operational metrics, which is why the authors recommend stratified, region-weighted searches and a simple cost–time–quality reporting template so future work tracks total cost of ownership alongside learning and equity outcomes.

Citation

Raza, S. H., & Farooq, A. (2025). Review of Artificial Intelligence in Education from 2020 to 2025. EdArXiv. doi:10.35542/osf.io/6bnez_v1.

Embed this page

Copy the code below to embed a chromeless version of this page in a learning management system or other website. The embedded view hides the site header, navigation, and footer.