Research Article
A systematic review of generative AI in education: Empirical insights from a human–AI interaction perspective
Synthesis: Liang, Yang, Sha, Gašević, Yan & Chen (2026) systematically review 56 empirical studies on GenAI in education through the AIED-HCD framework, analyzing three human–AI interaction modes along dimensions of human control and AI automation. They find that practice remains cautious toward high-AI-automation modes, but a high-control + high-automation mode is emerging as a trend — suggesting the future is not AI replacing humans but calibrated human–AI complementarity.
This BJET review synthesizes 56 empirical studies on GenAI in education, uniquely applying the AIED-HCD framework which conceptualizes human–AI interaction along two dimensions: human control and AI automation. Three interaction modes emerge: (1) low AI automation + high human control (teacher-led), (2) balanced, and (3) high automation + high human control (emerging trend). A sensitivity analysis validates the robustness of findings across modes. The review identifies that while practice remains cautious toward high-automation modes, the simultaneous presence of high human control with high AI automation represents a promising direction.
- 56 empirical studies systematically reviewed through AIED-HCD human–AI interaction framework
- Three interaction modes identified along human control × AI automation dimensions
- High-control + high-automation mode emerging as promising direction — not AI replacement but complementarity
- Most current practice remains in lower-automation modes with strong teacher/learner oversight
- Sensitivity analysis confirms findings robust across interaction modes
What this means for practice
- Instructors. Match the level of AI automation to the task instead of avoiding it. Most reviewed GenAI practice sits in low-automation modes with heavy teacher oversight, while the emerging high-control-plus-high-automation mode is the one the review identifies as the promising direction.
- Instructors. Keep humans on meaning-making and critical decision-making while letting GenAI execute structured instructional tasks; the externalization mode worked mainly for text generation and feedback-intensive tasks, and the review cautions against reading those gains as educational value.
- Researchers. Report assumption checks, effect sizes, and sample-size justification as a matter of course: only 12 of the 56 reviewed studies reported sample and effect sizes meeting the conventional power ≥ 0.80 threshold, and many did not document whether parametric assumptions were tested at all. Report the human–AI interaction mode alongside every result rather than aggregating across modes: the mode-specific sensitivity analysis is where the actionable differences between designs appear.
- Researchers. Prioritize hybrid intelligence designs with current models in controlled human–AI collaboration experiments. Only 9 of 56 studies used that mode and just 2 of those met the power criterion, so its promise currently rests on underpowered evidence.
- Faculty developers. Give teaching staff replicable templates, entry requirements, and evaluation tools for switching between interaction modes; the review's recommendation is a framework that specifies when to move from automation to more tightly coupled collaboration.
Limitations
- Search coverage. Four bibliographic databases (Web of Science, Scopus, ACM Digital Library, IEEE Xplore), English-language publications only, keyword selection choices, and a search run in December 2024 — the authors note relevant and rapidly emerging work may have been omitted.
- The evidence base is largely underpowered. Of 56 included studies, only 12 met the conventional power threshold, so mode-level conclusions about effectiveness rest on small studies — and on uneven subsets, since the most populated internalization mode alone accounts for n = 34 of them — while compliance rates are described as similar across modes rather than decisive.
- Conceptual boundaries and reported evidence only. The authors state the review still has limits in separating "learning" from "performance," and its statistical-rigour findings depend on what included studies reported, which was often incomplete.
Citation
Liang, Z., Yang, K., Sha, L., Gašević, D., Yan, L., & Chen, G. (2026). A systematic review of generative AI in education: Empirical insights from a human–AI interaction perspective.