acceptodds
Under review as a conference paper at ICLR 2027

Specialize to Protect: Comparative Advantage in Model–Harness Safety

Abstract

Modern large language model (LLM) systems rely on both model-internal safety behavior and external harnesses, yet it is unclear whether every layer should learn every defense. We study this question under matched training and deployment budgets. Controlled evidence and timing interventions show that models and harnesses develop complementary safety strengths because they observe different information and can intervene at different stages. Broad all-task training can weaken these strengths through resource dilution and cross-task interference. We therefore introduce a collaborative safety framework that assigns defenses to feasible interfaces, specializes model and harness components with retention constraints, and allocates further optimization using confirmed marginal end-to-end gains. Across 12 benchmark suites spanning harmful dialogue, tool-use agents, and secure code, with 7 trainable configurations from 3 model families and 6 additional frozen configurations, the framework improves agent safe-and-useful success by 6.1 percentage points over full-coverage training and by 1.2 points over the strongest matched-budget allocation baseline, with comparable benign-task performance. Information removal and restoration, assignment swaps, component ablations, held-out distributions, and adaptive attacks support the proposed mechanism. These results show that system safety can improve not by making every layer maximally general, but by exploiting comparative advantage across layers.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.