UniTrust: Unifying Risk Detection and Intervention with Minimal Task Supervision
Abstract
Adapting language models to new safety tasks remains costly, as doing so often requires large task-specific datasets for learning both what to detect and how to respond. We introduce UniTrust, a unified approach for adapting language models to new safety tasks from a compact task specification consisting only of a natural-language description and a handful of positive and negative examples. We formulate task-specific safety adaptation as complementary read and write operations over a shared hidden-representation space: task-relevant risk is read from frozen representations for prediction, while task-specific behavioral changes are written through activation interventions. Starting from minimal task-specific supervision, UniTrust trains a lightweight detector over frozen hidden representations to identify task-specific risks and derives activation directions from matched pairs of safe and unsafe responses to guide generation. To provide broader task coverage without additional human annotation, our Corpus-Grounded Synthesis procedure retrieves diverse contexts from public corpora and rewrites them according to the target task specification, constructing supervision for both detection and intervention. With only ten positive and ten negative real examples per task, UniTrust consistently improves detection across ten safety domains and three model families, and reduces average unsafe compliance by more than 50% across aligned models and their safety-weakened counterparts. Experiments demonstrate that effective safety adaptation can be built from a minimal task specification with little task-specific human supervision. Code and data are available at https://anonymous.4open.science/r/UniTrust.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.