Research Article
VETTING: A dual-LLM framework for in-loop safety verification via policy isolation in educational AI
Synthesis: VETTING: A dual-Large Language Models (LLMs) framework for in-loop safety verification via policy isolation in educational AI
Key Findings
- VETTING is a dual-LLM architecture that separates response generation from policy verification: a generator LLM produces responses and a separate verifier LLM checks them against safety policies at runtime.
- In an in situ deployment with 151 middle school students in a writing activity, VETTING achieved precision of .943, recall of .913 (95% CI [.652, .983]), and an F1 score of .928 for safety violation detection; human audit confirmed 50 true positives and 3 false positives, with an estimated 4.8 undetected violations.
- The deployment was associated with an estimated 91.2% reduction in inappropriate content exposure, at the cost of a 19.6% increase in token usage (628,011 tokens, roughly 12,560 tokens per prevented violation); median response latency rose from 2.9 s to 6.6 s on verification-triggered turns (Mann-Whitney U = 46,770.5, p < .001, Cohen's d = 0.66).
- Across 1250 student–AI interaction turns, only 51 turns (4.1%) triggered verification, concentrated in 25 conversations (16.6% of all conversations); the presence of risky keywords in prompts raised the odds of verification failure more than fourfold (OR = 4.59, 95% CI [2.04, 10.34], p < .001).
- The work documents a taxonomy of student boundary-testing behaviors observed during authentic classroom use: eight violation categories, led by Romantic or Intimate Relationship Themes (30.2% of instances) and Inappropriate or Mature Topics (26.4%), with students frequently rephrasing or escalating prompts after verification failures.
- Policy isolation keeps policy specifications hidden during interaction, making circumvention attempts monitorable and auditable rather than embedded in prompts.
- An Open Source Python implementation of the framework is available.
Study Design & Method
Educational AI systems increasingly rely on large language models to support student writing and inquiry, yet enforcing safety and instructional constraints during open-ended, multi-turn interaction remains challenging. Existing approaches commonly embed such constraints within conversational prompts or rely on static filtering; over time these approaches may become sensitive to user interaction, making it difficult to monitor and audit when students are able to circumvent or otherwise attempt to violate the measures. VETTING instead separates response generation from policy verification and applies explicit policy checks at runtime, illustrated through a grounded instantiation that enforces instructional and safety constraints without exposing policy specifications during interaction. The evaluation ran in a middle school classroom during a structured, timed writing activity: 190 students in grades 6–8 were given 45 minutes to write a 500-word essay on the advantages and disadvantages of AI in education, and 151 of them interacted with the chatbot. Every candidate response was checked by the verification layer before release; failed responses triggered an iterative rewrite loop bounded at three attempts before a fallback response was issued. Evaluation combined analysis of student–AI interaction behavior, human audit of verification outcomes against a thematic codebook, characterization of computational overhead, and a retrospective comparison with a single-LLM embedded-policy baseline.
What this means for practice
- Designers. Adopt policy-isolated runtime verification for student-facing systems instead of embedding safety rules in the system prompt: a retrospective comparison found that 35.3% of the violations VETTING intercepted would still have produced student-visible responses under embedded prompting, positioning separated verification as the higher-control design point for minors and developmentally sensitive content.
- Designers. Budget explicitly for the overhead, which is modest because verification fires rarely: an estimated 91.2% reduction in inappropriate content exposure cost a 19.6% increase in token usage (628,011 tokens, about 12,560 tokens per prevented violation) and raised median response latency from 2.9 s to 6.6 s on verification-triggered turns.
- Instructors. Prepare classroom guidance from the eight-category taxonomy of boundary-testing behaviors observed in authentic use — led by Romantic or Intimate Relationship Themes (30.2% of instances) and Inappropriate or Mature Topics (26.4%) — because students frequently rephrased or escalated prompts after verification failures.
- Designers. Keep policy specifications hidden during interaction so circumvention attempts stay monitorable and auditable, and treat risky keywords in prompts as a prioritization signal, since their presence raised the odds of verification failure more than fourfold (OR = 4.59, 95% CI [2.04, 10.34], p < .001).
- Administrators. Plan human oversight alongside automated checks: only 51 of 1250 student–AI interaction turns (4.1%), concentrated in 25 conversations (16.6%), triggered verification, so verifier-based safeguards in Human-in-the-Loop deployments complement rather than replace instructor review.
Limitations
- The taxonomy is exploratory rather than fully validated: the categories were developed through collaborative discussion and then applied by a single annotator, so formal inter-rater reliability metrics could not be computed.
- Recall estimates come from a sample-based audit of interactions that passed the guardrail, so the confidence intervals are wide because violations have a low base rate.
- The evaluation sits inside one middle school writing activity centered on AI in education, which may shape both what students typed and which violations appeared.
- It was not designed as a direct empirical comparison against strengthened prompt-based safeguards, and the retrospective baseline does not reproduce live-interaction dynamics.
Citation
Li, H., Zhang, S., & Botelho, A. F. (2026). VETTING: A dual-LLM framework for in-loop safety verification via policy isolation in educational AI.