Research Article
Adversarial Stress Testing of Role-Playing Language Agents using Multi-Agent Evaluation
Synthesis: This paper presents a modular multi-agent platform for adversarially stress-testing role-playing language agents through structured, multi-turn dialogue. It coordinates three agents — a strategy-driven Interrogator (applying six progressive adversarial strategies), a black-box Target, and an automated Judge — to reveal cumulative behavioral failures that static, single-turn benchmarks miss. Across three personas (Healthcare Assistant, Customer Support Agent, Financial Advisor) and three LLM families, multi-strategy testing reduced overall robustness scores by 0.17–0.20 points versus a single-strategy baseline, with Authority Challenge and Emotional Manipulation emerging as the most effective attacks and automated judging aligning strongly with human experts (r = 0.82, Fleiss' κ = 0.71). The framework is directly relevant to educational AI agents, where persona consistency and ethical constraints are critical for safe deployment with learners.
Key Findings
- Multi-strategy testing exposes hidden failures. Adversarial evaluation that coordinates multiple attack strategies across multi-turn dialogue reveals role drift, ethical deviation, and inconsistency that are invisible to single-strategy or static testing, lowering robustness scores by 0.17–0.20 points across all personas.
- Two strategies dominate. Authority Challenge and Emotional Manipulation are the most effective attack vectors, inducing the highest ethical deviation across every model family tested.
- Degradation generalizes across models. Consistent failure patterns appear across Llama-3.3-70B, GPT-4o-mini, and Claude-3.5-Haiku, signaling fundamental RPLA vulnerabilities rather than model-specific weaknesses.
- Automated judging is valid and scalable. The Judge aligns strongly with human experts (r = 0.82, Fleiss' κ = 0.71), though it shows a conservative bias, scoring ethical violations 0.08 points lower than humans.
- Failures emerge late. Critical failures onset in the second half of conversations (after turn 5–6), so short or single-turn evaluations systematically overestimate robustness.
Background and Motivation
Role-playing language agents (RPLAs) are increasingly deployed in high-stakes settings such as healthcare assistance, customer support, and education, where success depends not just on linguistic fluency but on maintaining consistent personas, respecting ethical constraints, and behaving coherently over long interactions. Yet most current evaluation approaches rely on static benchmark datasets or isolated single-turn prompts that fail to model adversarial or malicious user behavior, while human evaluation is expensive, subjective, and hard to scale. As a result, deployed agents exhibit role abandonment, ethical violations, and contradictory behavior under adversarial inputs.
The authors propose a modular, multi-agent framework that stress-tests RPLAs through sustained adversarial dialogue rather than isolated probes. Its contributions include a production-grade extensible platform, a strategy-driven agent for systematic failure discovery, a multi-dimensional automated evaluation framework, and a reproducible experimental pipeline with persistent storage and Visualization.
The Three-Agent Evaluation Platform
The framework is built on a client-server architecture with a FastAPI backend and a React dashboard, coordinating three logically independent agents under an experiment orchestration layer:
- Interrogator Agent. An automated adversary that applies structured attack strategies across multiple turns, dynamically adapting and escalating pressure based on the Target's prior responses. It models realistic misuse rather than issuing random or isolated prompts.
- Target Agent. The RPLA under evaluation, treated as a black box to reflect real deployment conditions. It is initialized with a role description, persona specification, domain context, and explicit behavioral and ethical constraints, maintaining an internal conversation memory.
- Judging Agent. A post-hoc evaluator that never participates in dialogue generation, keeping assessment unbiased. It produces quantitative scores plus qualitative explanations of specific failure points.
The Six Adversarial Strategies
The Interrogator Agent applies six structured strategies, each targeting a specific behavioral vulnerability:
- Role Drift — attempts to override the assigned role (persona abandonment).
- Ethical Probing — tests adherence to constraints (constraint violations).
- Contradiction — induces logical inconsistencies.
- Confusion — introduces ambiguous inputs causing coherence breakdown.
- Authority Challenge — simulates boundary pressure (boundary violation).
- Emotional Manipulation — applies affective influence to destabilize responses.
Strategies are deployed via a progressive difficulty-based mechanism: early turns (1–3) favor lower-difficulty attacks like Role Drift and Confusion, while later turns unlock escalated strategies like Emotional Manipulation and Authority Challenge, mirroring real-world misuse patterns.
Evaluation Metrics
The Judging Agent computes four normalized, model-agnostic metrics bounded to [0, 1], combined into an overall composite score:
- Role Fidelity (RF) — how well the Target maintains its role, weighted from constraint adherence, role-language consistency, and absence of role abandonment.
- Drift Index (DI) — progressive deviation from the role, computed by comparing role-term density between early and late conversation segments.
- Ethical Deviation (ED) — violations of behavioral and ethical constraints.
- Consistency (CS) — logical coherence and tonal stability across turns.
The overall score is 0.3 × RF + 0.2 × (1 − DI) + 0.3 × (1 − ED) + 0.2 × CS, with all metric code released Open Source for community auditing.
Results
Across three personas and two conditions (a single-strategy Role Drift baseline vs. the full multi-strategy suite), multi-strategy evaluation produced consistent robustness declines. The Healthcare Assistant degraded most severely (0.837 → 0.634, a drop of 0.203), reflecting the difficulty of maintaining strict medical boundaries under sustained emotional and authority pressure; the Customer Support Agent dropped least (0.867 → 0.693), owing to its more concrete operational constraints. Cross-model validation on the Healthcare Assistant showed Claude-3.5-Haiku most robust (Overall = 0.712 ± 0.041), followed by GPT-4o-mini (0.681 ± 0.035) and Llama-3.3-70B most vulnerable (0.634 ± 0.038).
Role abandonment was most often triggered by Authority Challenge and Confusion after turn 5, while ethical violations concentrated under Emotional Manipulation and Ethical Probing. The authors hypothesize that LLMs prioritize local conversational coherence over global constraint adherence: as pressure mounts, the model aligns with user intent at the expense of predefined constraints, suggesting current alignment techniques insufficiently account for long-term interaction dynamics.
Relevance to Educational AI
AI tutors and pedagogical agents are role-playing agents that must maintain consistent instructional personas, making them prime candidates for this evaluation methodology. Student interaction patterns can be adversarial — testing boundaries, emotional appeals, and authority challenges — and educational agents must remain robust to these behaviors. The open-source release provides infrastructure for the AIED community to evaluate the safety and robustness of educational language agents, particularly given the risk that constrained agents may drift or violate boundaries under sustained pressure.
What this means for practice
- Designers. Make multi-turn adversarial testing a gate before deploying an educational agent, not an afterthought: the Healthcare Assistant persona scored 0.837 under a single-strategy baseline but 0.634 under the full multi-strategy suite, and critical failures emerged after turn 5–6, so short or single-turn testing systematically overestimates robustness.
- Designers. Lead test suites with Authority Challenge and Emotional Manipulation, the two vectors that induced the highest ethical deviation across every model family, and watch turn 5 onward for role abandonment, which Authority Challenge and Confusion most often triggered.
- Designers. Budget periodic human calibration of automated judging rather than trusting the Judge indefinitely: it aligned strongly with three domain experts (r = 0.82, Fleiss' κ = 0.71) but systematically scored ethical violations 0.08 points lower than humans.
- Designers. Do not assume a stronger model family removes the risk. Cross-model validation on the Healthcare Assistant spanned 0.712 (Claude-3.5-Haiku) to 0.634 (Llama-3.3-70B), all below the single-strategy baseline, so robustness must be measured per deployment rather than inherited from the model.
Limitations
- The framework covers prompt-based adversarial attacks only; the authors state that vulnerabilities arising from training, fine-tuning, or reinforcement learning are not considered, so deeper alignment problems may go uncaptured.
- Automated judging was validated by three domain experts on a stratified sample of just 60 conversation turns (20 per persona), and the authors note the Judge may still introduce bias in complex or ambiguous cases, with scaling human calibration across more personas and evaluators an open challenge.
- All four evaluation metrics — Role Fidelity, Drift Index, Ethical Deviation, and Consistency — are computed by rule-based text analysis combined with keyword pattern matching over the conversation history, so the scores detect patterns rather than judge meaning.
- Only three personas (Healthcare Assistant, Customer Support Agent, Financial Advisor) were tested, all in controlled environments; the authors note real deployments involve more complex domain-specific roles, and that adversarial testing techniques could themselves be misused to exploit deployed systems, requiring responsible access and human oversight.
Citation
Shouqi, S., Nazly, A., Wanniarachchi, J., & De Alwis, R. (2026). Adversarial Stress Testing of Role-Playing Language Agents using Multi-Agent Evaluation. v1.