Research Article
CyberAGENTS: Structured Autonomy for Agentic Gamified Learning in Cybersecurity
Synthesis: Hornung et al. (2026) present CyberAGENTS, an agentic framework for gamified cybersecurity learning that enables structured autonomy through ontology-guided validation, schema-governed behavioral control, and competency-based progression. The learning loop is decomposed into four specialized agents — challenge, support, evaluation, reward — each governed by behavioral schemas that bound autonomy without eliminating generative flexibility, while a cybersecurity ontology validates all generated content before display. Classroom deployment with 24 undergraduates, complemented by expert evaluation and an Large Language Models (LLMs)-as-judge ablation, found improved engagement, clearer Feedback interpretation, and greater learner Trust in AI-generated responses when behavioral schemas and ontology validation are active. The work positions structured autonomy as a design and ethical principle for safe, pedagogically aligned agentic education systems, and connects to CS Education, Game-Based Learning, and Generative AI themes.
Background: Gamified Cybersecurity Learning
Gamification is especially effective in learning domains requiring active Problem Solving and iterative skill-building, such as cybersecurity education. Through progressive challenges, adaptive instructional support, structured feedback loops, and reward progression, gamified design enhances cognitive engagement and retention, particularly in practice-driven domains. Active learning matters here because cybersecurity concepts are abstract, adversarial, and high-stakes, and traditional lecture-driven instruction rarely supplies iterative reasoning practice or timely corrective feedback.
Generative AI agents offer a path to delivering such gamified experiences adaptively at scale, but they introduce well-documented risks in educational settings: inconsistent behavior, hallucinated reasoning, and misalignment with pedagogical frameworks. Unconstrained multi-agent systems can produce conflicting feedback or pedagogically misaligned behavior, and dynamic prompting without explicit control can lead to instructional drift. Grounding these systems in learning science is therefore essential. CyberAGENTS addresses the central tension — bounding agentic autonomy while preserving generative flexibility — through three layered control components: competency-based progression, schema-governed behavioral control, and ontology-guided validation.
Competency-Based Progression
CyberAGENTS begins by extracting structured competencies from course materials using a cybersecurity-specific competency schema. The extractor identifies key entities such as tools, techniques, concepts, attacks, vulnerabilities, and defenses; assigns difficulty levels from Beginner to Expert; and maps instructional relations relevant to gameplay — prerequisites, concept clusters, and learning enablers. The resulting competency graph encodes topical coverage and progression structure, serving as the initialization blueprint for challenge generation.
This competency-based progression keeps gameplay curriculum-aligned from the outset. Rather than sampling arbitrary prompts, the system draws on a leveled representation of domain knowledge, letting challenge progression follow explicit instructional structure. This is especially important in cybersecurity education, where conceptual dependencies and difficulty progression must be carefully managed to avoid cognitive overload or fragmented learning. In this way the design operationalizes scaffolded instruction, sequencing tasks by prior knowledge and prerequisite relationships.
The Agentic Learning Loop
Gamification in CyberAGENTS is operationalized through a structured multi-agent learning loop. Instead of a single generative model improvising across multiple instructional functions, the framework decomposes gameplay into four specialized pedagogical agents: ChallengeAgent, BuddyAgent, CriticAgent, and RewardAgent. Together they enact the cycle of challenge → learner attempt → evaluation → progression.
At runtime, a central orchestrator coordinates these agents and maintains shared state across turns, including topic, difficulty, hint usage, feedback history, and XP trajectory. This shared state preserves coherent progression while adapting support and challenge difficulty to learner behavior:
- The ChallengeAgent generates competency-aligned tasks with calibrated difficulty and explicit objectives.
- The BuddyAgent provides adaptive instructional support based on learner performance and interaction signals, offering hints and guided options.
- The CriticAgent evaluates learner responses and produces formative feedback anchored to explicit criteria.
- The RewardAgent assigns XP, badges, and progression indicators according to a mastery-oriented reward rubric tied to motivational reinforcement.
This role-based decomposition supports the core gamified learning functions of progressive challenges, adaptive support, structured feedback, and mastery-based rewards. More importantly, it enables modular control over instructional dynamics, making the loop more interpretable and stable than monolithic prompting approaches — a more transparent tutoring architecture.
Schema-Governed Behavioral Control
To keep adaptive agent behavior instructionally coherent, CyberAGENTS defines explicit behavioral schemas for each agent. These schemas function as structured metadata contracts specifying how an agent operates, when it intervenes, and what form its outputs may take, parameterized through fields governing operational mode, difficulty level, trigger conditions, response style, evaluation criteria, and reward logic.
These schemas implement structured autonomy: agents retain generative flexibility within defined bounds but cannot deviate from prescribed instructional roles. Challenge generation is constrained by difficulty and explicit objectives; instructional support is regulated through trigger conditions and assistance levels that enable gradual fading of help as competence increases — preventing over-assistance; evaluation follows structured criteria; and rewards are tied to an experience-based mastery rubric. This prevents role drift while preserving adaptivity across learners, and it supports cross-agent coordination because agents expose structured properties (difficulty, hint usage, grading mode, XP tier) that let the orchestrator maintain coherent progression across turns.
Ontology-Guided Validation
While schemas regulate instructional behavior, domain correctness is enforced through ontology-guided validation. CyberAGENTS integrates AISecKG, a cybersecurity ontology that defines valid entity types (tool, technique, attack, vulnerability, defense, system) and permissible relations among them (exploits, detects, counters, uses, can harm). This ontology acts as a reasoning constraint layer applied to agent-generated content before learner exposure.
During gameplay, outputs from the ChallengeAgent, BuddyAgent, and CriticAgent are checked against ontology-defined entity categories, relation patterns, and unsafe-content rules. Content that violates semantic constraints or safety thresholds is flagged and re-prompted under stricter conditions before presentation. This pre-display gating complements schema-governed behavioral control: schemas constrain pedagogical behavior and progression logic, while ontology validation constrains cybersecurity reasoning and semantic validity. Together they regulate both instructional conduct and domain semantics — reducing hallucinations and unsafe instructional drift in a dynamically generated learning environment, and exemplifying Pedagogical Safety as a design constraint.
Evaluation
CyberAGENTS was implemented as a web-based interactive system (React frontend, Python Flask backend, deployed on Google Cloud Run) with LLM inference served via the Together AI API using Llama-based models. Participants accessed it through a public web portal and completed a short novice-level session of cybersecurity challenges, then a post-study survey.
Human study. A within-subjects design with 24 undergraduate students, plus expert feedback from cybersecurity and educational-design specialists, assessed engagement, feedback quality, trust in AI guidance, and perceived learning value. Quantitative results were positive across most dimensions, with scenario authenticity highest (M = 4.21), followed by gamification impact, trust in system feedback, and willingness to recommend. Lower scores on challenge difficulty alignment (M = 3.12) suggested room for improvement in challenge balancing; Cronbach's alpha was α = 0.90. Qualitative feedback surfaced four themes — trust in AI scaffolding, clarity of feedback, engagement and interaction design, and task authenticity — with participants valuing the BuddyAgent's guided hints while requesting shorter, more structured responses. This combination of survey, observational, and open-ended evidence reflects a mixed-methods evaluation approach.
LLM-as-judge evaluation and ablation. An LLM-as-judge assessed full transcripts on learner intent alignment, challenge quality, interaction coherence and state tracking, agent role fidelity, and domain relevance/cybersecurity grounding. The full system scored highest on domain relevance and cybersecurity grounding (4.46) and agent role fidelity (4.35). An ablated version with the expert-informed structural components (competency map and ontology constraints) removed scored lower on all dimensions, with the largest degradations in challenge quality, learner intent alignment, and domain grounding — evidence that the structural components matter. The authors interpret these ablation results descriptively, since the ablation deployment received less traffic.
Key Findings
- CyberAGENTS grounds gamified cybersecurity learning in competency-based progression, organizing topics by difficulty and prerequisite relationships so gameplay is curriculum-aligned and scaffolded rather than sampled arbitrarily.
- Decomposing the learning loop into four schema-governed agents (challenge, support, evaluation, reward) enacts structured autonomy, bounding agent behavior and preventing role drift while preserving generative flexibility.
- Ontology-guided validation with the AISecKG cybersecurity ontology gates all agent outputs before display, enforcing domain-consistent reasoning and safety constraints that reduce hallucinated or unsafe content.
- A human study with 24 undergraduates and expert evaluators found improved engagement, clearer Feedback interpretation, and greater learner trust in AI responses when behavioral schemas and ontology validation were active.
- An LLM-as-judge ablation showed the full framework outperforming an unconstrained configuration, especially on challenge quality, learner intent alignment, and domain grounding, supporting the role of structured control in stabilizing instructional behavior.
What this means for practice
- Instructional designers. Bound generative flexibility through structured autonomy rather than scripting it away or leaving it unconstrained: layer behavioral schemas that regulate instructional conduct over ontology validation that constrains domain reasoning, so agentic learning systems stay in role and domain-aware in high-stakes technical domains.
- Instructional designers. Anchor evaluation and Feedback to explicit schema-defined criteria and return short, structured responses, so learners can interpret and act on formative feedback instead of trusting an opaque generative judgment.
- Edtech designers. Decompose the learning loop into specialized, schema-governed agents (challenge, support, evaluation, reward) rather than letting one model improvise across every instructional function, which keeps instructional dynamics interpretable and stable.
- Edtech designers. Treat pre-display ontology validation and schema-bounded behavior as safeguards requiring continued human oversight, because cybersecurity is adversarial and a learner may act on generated procedural guidance.
- Researchers. Reuse the mixed-methods plus LLM-as-judge ablation design to test whether the structural components, rather than the model, produce the gains, since the ablation here was run only descriptively.
Limitations
- The human study rests on 24 undergraduates in a single within-subjects session of roughly 5–15 minutes with no control condition, so the positive ratings describe one short novice-level interaction rather than a learning effect.
- Every learner measure is a self-reported post-study survey item (plus one open-ended question); the item means range from scenario authenticity at 4.21 down to challenge difficulty alignment at 3.12, and perceptions are not a measure of learning or retention.
- The LLM-as-judge results are single-model judgments of transcripts, and the ablated configuration received less traffic than the full system, so the authors read the ablation descriptively rather than as a controlled comparison.
- Evaluation used one novice difficulty level and Llama-based models served through a single commercial API, so performance may not transfer to other models, difficulty levels, or longer curricula.
Citation
Hornung, I., Marasinghe Arachchige, D., Kumarage, T., Agrawal, G., Deng, Y., Chen, Y.-C., & Liu, H. (2026). CyberAGENTS: Structured autonomy for agentic gamified learning in cybersecurity.