π Research Article
Child Safety in Generative AI: An Expert-Guided and Incident-Grounded Evaluation Framework
Haein Kong β HEAL Workshop at CHI 2026, submitted 1 Jul 2026
Haein Kong β HEAL Workshop at CHI 2026, submitted 1 Jul 2026
Proposes an evaluation framework for child-specific harms in generative AI; applied to education domain, Llama Guard models struggle to detect unsafe user prompts from children.
Key Findings
Study Design & Method
The framework combines two evidence sources: hazard categories derived from expert guidelines and categories mined from AI incident databases. These inform a synthetic test set in which harmful and safe education-domain user prompts are generated from incident descriptions, with the user assumed to be a teen or student. The resulting test set is used to evaluate safety classifiers β here, three Llama Guard models β on their detection of unsafe user prompts, with assessments scored as safe or unsafe. This design lets the authors measure child-specific safety performance in a region where existing general-population benchmarks leave a gap.
Implications for AI in Education
The results carry a direct warning for AI-based learning environments: general-purpose safety classifiers do not reliably catch education-related unsafe prompts from children, so Pedagogical Safety cannot be assumed from standard model safeguards. Schools and edtech providers deploying Generative AI tools need child-specific evaluation, incident-grounded testing, and human oversight rather than reliance on off-the-shelf safety models. The framework's structure β expert guidance plus incident data plus synthetic testing β is itself a template that educational institutions and researchers can reuse to evaluate tools for younger users, with implications for Privacy and Equity in who is protected by default safety practices.
Connected Concepts
Connected Articles
Citation
Haein Kong (2026). Child Safety in Generative AI: An Expert-Guided and Incident-Grounded Evaluation Framework. arXiv:2607.00395. HEAL Workshop at CHI 2026, submitted 1 Jul 2026