acceptodds
Under review as a conference paper at ICLR 2027

Guard-to-Action: Multilingual Long-Context Safety in Tool-Using Agents

Abstract

Large language model (LLM) agents are increasingly deployed with access to real-world tools, yet most safety evaluations remain English-only and short-context, leaving open whether multilingual translation and long-context placement can jointly weaken input guardrails and downstream agent behavior. We introduce a large-scale multilingual, long-context safety benchmark built from 1,901 unsafe English root prompts spanning 19 harm categories (derived from AEGIS2.0), translated into 98 languages to yield up to 186,298 prompt instances with translation-quality metadata. Each translated prompt is screened by five runtime input guards (AprielGuard, CREST, GuardReasoner, WildGuard, XGuard); guard-admitted prompts are then embedded, unchanged, in benign English conversation history at 8K and 32K token lengths and at the beginning, middle, or end of the context, and forwarded to tool-using downstream agents that must refuse, escalate, or select a restricted (inert) tool. Using matched, sample-level comparisons with paired-bootstrap confidence intervals and exact statistical tests, we find that downstream safety depends jointly on agent identity, context length, and prompt position, with no single condition uniformly safest, and that increasing context length can raise or lower risk depending on this interaction rather than acting uniformly in either direction. We further find that low-resource languages exhibit substantially higher downstream risk than high- or medium-resource languages despite comparable guard admission rates, and that aggregate safety rates across context lengths can mask considerable prompt-level instability, with a nontrivial share of individually matched prompts flipping between safe and unsafe outcomes. Our current evaluation covers two open-weight instruction-tuned agents (Llama-3.1-8B-Instruct and Qwen2.5-14B-Instruct); we are extending this evaluation to additional agent families to assess how broadly these interaction effects generalize.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.