Research Article
Artificial intelligence in vocational education and training: A systematic review of educational purposes, theoretical conceptualizations, and empirical effectiveness
Synthesis: First systematic review of AI in vocational education and training, identifying 26 empirical studies (2015–2026) via ERIC, Web of Science, and Elicit, analyzed with a theory-informed coding scheme under PRISMA guidelines.
Key Findings
- This is the first systematic review of AI in vocational education and training (VET), identifying 26 empirical studies published between 2015 and 2026 through ERIC, Web of Science, and Elicit, analyzed with a theory-informed coding scheme following PRISMA guidelines.
- The corpus spans 9 technical-domain studies, 3 in health, 5 in business administration and services, and 9 domain-general studies; settings were 6 classroom, 8 online, 4 blended, and 8 simulation-based — with no study conducted in workplace settings despite VET's work-based character.
- Research is geographically concentrated: 17 of 26 studies originated in Asia (China, South Korea, Singapore, Taiwan, Thailand, Indonesia), with the remainder from European contexts (Germany, the Netherlands, Norway) and isolated contributions (New Zealand, Saudi Arabia, Turkey); the field is also fragmented, as only 9 studies shared at least one reference and none cited each other directly.
- Intelligent Extended Reality (XR) shows consistent positive effects on procedural competence, practical skills, and learner motivation — e.g., a randomized pre-post comparison (iXR n = 14 vs. traditional group task n = 15) found both groups gained knowledge but gains were significantly higher with iXR.
- Intelligent Tutoring Systems foster declarative and procedural knowledge, while AI chatbots show promising effects on self-AI Regulation in Education and task performance — including a grounded-theory study of 408 polytechnic students tracing a self-regulatory arc (goal setting → performative interaction → reflection), though the only randomized chatbot trial (n = 50, posttest-only) reported advantages on seven competencies without effect sizes, pretests, or baseline checks.
- The evidence base is methodologically constrained: only five randomized experimental studies were identified among the 26, 21 of 26 rely on pre-experimental or quasi-experimental designs, and most measure outcomes immediately after the intervention; affective and meta-cognitive outcomes rest predominantly on self-report, raising novelty-effect concerns.
- Only three studies explored AI-empowered designs that grant learners an active role; meta-cognitive goals such as self-regulated learning are frequently espoused but rarely implemented through genuinely learner-empowered systems.
- Across applications, AI is predominantly implemented through behaviorist or cognitively oriented instructional designs that emphasize drill-and-practice and adaptive feedback, while approaches fostering learner agency, critical reflection, and autonomous decision-making remain underrepresented.
- Current research largely reflects a generalized "success narrative"; the authors call for future studies of failure cases, contextual moderators, and boundary conditions to develop a more differentiated understanding of effectiveness.
Study Design & Method
This is a PRISMA-guided systematic review of 26 empirical studies (2015–2026) identified through ERIC, Web of Science, and Elicit, with Scopus added as a supplementary domain-specific database, and coded with a theory-informed scheme that distinguishes AI-directed, AI-supported, and AI-empowered human-AI interaction paradigms and underlying learning theories; all included studies were independently double-coded, with coding documented in a publicly available dataset (Appendix 2). Methodologically, quantitative designs dominate (14 studies, surveys most frequent), followed by mixed methods (10) and two qualitative case studies; sample sizes range from 9–15 VET learners (qualitative) to 20–3,518 (quantitative). Among the quantitative studies, 4 employ quasi-experimental approaches and 4 use randomized experiments per the methodological breakdown, and 2 rely exclusively on self-report while 5 use only objective measures (performance tests or log data).
The Constructivist Paradox and the Turing Trap
The review documents a notable paradox: constructivist theories are espoused in VET discourse while behaviorist AI implementations dominate in practice. The authors warn against an educational "Turing Trap" — the danger of using AI to replicate human instruction rather than to augment human judgment. Realizing the transformative potential of AI in VET, they argue, requires learning environments that augment human judgment, strengthen learner agency, and support teachers, rather than systems that merely automate existing instructional patterns.
What this means for practice
- Researchers. Study AI where VET actually happens — none of the 26 included studies was conducted in a workplace setting despite VET's work-based character. Add delayed post-tests, objective performance measures, and checks of transfer to job tasks instead of immediate self-report after the intervention.
- Administrators. Demand stronger causal designs before scaling: only 5 of the 26 studies were randomized experiments, while 21 relied on pre-experimental or quasi-experimental designs.
- Instructors. Match the tool to the learning goal the corpus actually supports — intelligent XR for procedural and practical skills, intelligent tutoring systems for declarative and procedural knowledge, and chatbots for self-regulation support — and prioritize Simulation and authentic, practice-proximal environments.
- Designers. Build systems that grant learners an active role rather than automating existing instruction. Only 3 of the 26 studies supported AI-empowered designs, and the authors warn of an educational "Turing Trap" when AI replicates a teacher instead of augmenting learner agency and human judgment.
- Researchers. Report failure cases, contextual moderators, and boundary conditions; the current corpus is dominated by a generalized success narrative that makes differentiated conclusions about effectiveness impossible.
Limitations
- The review restricted its search to English-language, peer-reviewed journal articles, likely excluding gray literature and non-English work — a notable gap given the applied, project-based and often locally documented nature of VET interventions.
- The database-dependent search may underrepresent regions with distinct publication cultures, partly explaining the Asia/Europe concentration.
- The use of Elicit as an AI-assisted discovery tool constrains full reproducibility, because retrieved outcomes depend on probabilistic ranking mechanisms and database coverage changes.
- The heterogeneity of included studies and frequent lack of transparency about AI implementations and instructional designs required interpretive judgment in coding, despite double-coding and consensus-based resolution of discrepancies.
Citation
Deutscher, V., Thomann, H., Zlatkin-Troitschanskaia, O., Weyland, U., Abele, S., Danek, A. H., Greiff, S., Rausch, A., Seeber, S., Seifried, J., & Winther, E. (2026). Artificial intelligence in vocational education and training: A systematic review of educational purposes, theoretical conceptualizations, and empirical effectiveness.