acceptodds
Under review as a conference paper at ICLR 2027

Can Agents Safely Expand Their Capabilities? An Empirical Study of Trust Decisions in Autonomous Capability Discovery

Abstract

As Large Language Model (LLM) agents seek capabilities beyond their predefined toolsets, they must decide which third-party software to trust and adopt, creating software supply chain risk during capability discovery. We study this adoption process from search and inspection through recommendation, installation, and execution. Our main study comprises 3,216 trajectories from six LLM agents across eight capability domains. We combine discovery over a frozen corpus of benchmark labeled clean tools and real GitHub repositories with a Resolution IV factorial experiment manipulating five trust signals under recommendation only and enabled execution conditions. Agents typically search once and inspect one candidate, surfacing only 34.6%-44.6% of the benchmark clean alternatives. In controlled trials, poisoned recommendation rates vary substantially across models, from 0% to 73.4%. Unverified security badges increase poisoned selection in multiple models, including by 28.1 percentage points for Qwen3.8-Flash. Under Execution authority, unsafe execution rates range from 0% to 62.9%. Several agents exhibit lower risk alongside greater abstention and reduced protocol completion. Additional guidance improves clean recommendation relative to a generic warning, yet both warnings increase abstention and reduce clean execution from baseline.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.