acceptodds
Under review as a conference paper at ICLR 2027

CONTRA: Red-Teaming Configurations of Personalizable Agents

Abstract

Recent tools such as OpenClaw have extended the capabilities of LLM-based agents from simple dialog-based systems to fully autonomous agents. These systems allow personalization through modifiable internal files and the installation of skills. While this enables deployment in a wide range of settings and the automation of diverse tasks, greater capability and autonomy increase the risk of harmful actions being executed unintentionally. In this work, we explore the interplay between agent configuration and the risk of executing dangerous actions without explicit instruction. To this end, we propose CONfiguration Tree-search for Red-teaming Agents (CONTRA), an LLM-assisted tree-search algorithm that discovers agent configurations resulting in the execution of harmful actions. CONTRA works by reasoning about dangerous configurations and evaluating them in a simulated environment. We construct SkillSafetyBench, a dataset of the 532 most downloaded skills from a public repository, along with 1869 harmful actions. In a large-scale analysis, we find that 74.2% of skills have at least one configuration resulting in the execution of a harmful action, which is substantially more than existing attacks. Utilizing a virtual machine, we further confirm that these vulnerabilities persist under real tool execution. Finally, we demonstrate that the discovery of successful configurations can be used to improve the safety of skills by leveraging them to implement safety features.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.