Synthesis: This paper presents a modular multi-agent platform for adversarially stress-testing role-playing language agents through structured multi-turn dialogue. With three coordinated agents â Interrogator (applying six progressive adversarial strategies), Target, and Judge â the system reveals failure modes invisible to single-strategy testing, reducing robustness scores by 0.17-0.20 points. The framework is directly relevant to educational AI agents, where persona consistency and ethical constraints are critical for safe deployment with learners.
Platform Architecture
The evaluation platform coordinates three specialized agents:
1. Interrogator Agent:
Applies six progressive adversarial strategies including Authority Challenge and Emotional ManipulationEscalates pressure across multi-turn interactions to test cumulative robustness2. Target Agent:
The role-playing language agent (RPLA) under evaluationEvaluated across diverse personas including educational tutoring roles3. Judging Agent:
Automated scoring across four dimensions: role fidelity, behavioral drift, ethical deviation, and consistencyAchieves strong human alignment (r = 0.82, Fleiss' Îș = 0.71)Key Findings
| Finding | Result |
|---|
| Multi-strategy vs. single-strategy | 0.17-0.20 point robustness reduction |
| Most effective attack | Authority Challenge + Emotional Manipulation |
| Cross-model consistency | Consistent degradation across Llama-3.3-70B, GPT-4o-mini, Claude-3.5-Haiku |
| Automated judging quality | r = 0.82 correlation with human judges |
Failure mode discovery: Multi-strategy adversarial testing reveals behavioral failures invisible to standard single-turn benchmarksStrategy effectiveness: Authority Challenge and Emotional Manipulation emerge as the most effective attack vectorsCross-model validation: Degradation patterns are consistent across three major LLM families, suggesting fundamental vulnerabilities rather than model-specific weaknessesRelevance to Educational AI
The framework's relevance to education is twofold:
AI tutors and pedagogical agents are role-playing agents that must maintain consistent instructional personas, making them candidates for this evaluation methodologyStudent interaction patterns can be adversarial (testing boundaries, emotional appeals, authority challenges), and educational agents must be robust to these behaviorsThe open-source release provides infrastructure for the AIED community to evaluate safety and robustness of educational language agentsConnected Concepts
Agentic AIAI EducationPedagogical SafetyAI TutoringPedagogical AgentConnected Articles
Detecting LLM Generated Text Latent Prompt â Once a Response, Always a Response: Detecting LLM-generated Text via Latent Prompt RestorationAgentic AI Education Scoping Review â Agentic AI in Education: A Scoping Review of Research Landscape, Capabilities, and the Frontier Agent ParadigmJeon Isd Agent Bench 2026 â ISD Agent BenchmarkMOOC To Maic â From MOOC to MAIC: Reshaping Online Teaching and Learning through LLM-driven AgentsEduagentbench Agent Teaching Benchmark â Are Agents Ready to Teach? A Multi-Stage Benchmark for Real-World Teaching WorkflowsAgentic Workflows Education â Agentic Workflows in EducationCitation
Shouqi, S., Nazly, A., Wanniarachchi, J., & De Alwis, R. (2026). Adversarial Stress Testing of Role-Playing Language Agents using Multi-Agent Evaluation. arXiv:2608.03166v1.