🏷️ safety
3 pages tagged with safety(3 articles, 0 concepts)
📄 ELBench: A Multi-Dimensional Benchmark for Education-Facing Large Language Models
> **Synthesis:** Jiang et al. (2026) introduce **ELBench**, the first benchmark to evaluate education-facing LLMs on all four required dimensions — General Capability, Safety and Trustworthiness, Basi…
📄 Adversarial Stress Testing of Role-Playing Language Agents using Multi-Agent Evaluation
> **Synthesis:** This paper presents a modular multi-agent platform for adversarially stress-testing [[agentic-ai|role-playing language agents]] through structured multi-turn dialogue. With three coor…
📄 Once a Response, Always a Response: Detecting LLM-generated Text via Latent Prompt Restoration
> **Synthesis:** EchoPrompt introduces a training-free zero-shot detector for [[plagiarism-detection|LLM-generated text]] that exploits the latent prompt dependency inherent in machine-generated conte…