acceptodds
Under review as a conference paper at ICLR 2027

SaplingGuard: A Multidimensional-Profile-Aware Multi-Agent Guardrail for Developmentally Safe Adolescent–LLM Interaction

Abstract

As adolescents increasingly use LLMs in everyday life, ensuring safe and developmentally appropriate responses has become essential. However, existing LLM guardrails primarily target explicit harmful content in isolated prompts or responses and are less effective at identifying implicit, context-dependent developmental risks. To address this limitation, we propose SaplingGuard, a plug-and-play, profile-aware and dialogue-aware guardrail that requires no modification to downstream model parameters. SaplingGuard decomposes adolescent safety intervention into three specialized agents for user profile construction, context-aware risk assessment, and intent-preserving prompt optimization. Together, these agents leverage the current prompt, preceding dialogue, and structured user characteristics to identify contextual risks and guide downstream response generation. We evaluate SaplingGuard on SaplingBench, which contains 276 three-turn dialogues spanning seven categories of developmental risk. Across ten adolescent profile conditions and nine open- and closed-source downstream LLMs, profile-aware retrieval improves the Major Hit rate from to . End-to-end intervention further reduces the average harmful response rate from to and increases the average safety score from to . These results show that user-profile and dialogue context provide complementary signals for identifying implicit developmental risks, and that SaplingGuard can serve as an effective external safety layer for adolescent–LLM interaction.Our code and benchmark are available at https://anonymous.4open.science/r/Sapling-1A07/.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.