acceptodds
Under review as a conference paper at ICLR 2027

UniTrust: Unifying Risk Detection and Intervention with Minimal Task Supervision

Abstract

Adapting language models to new safety tasks remains costly, as doing so often requires large task-specific datasets for learning both what to detect and how to respond. We introduce UniTrust, a unified approach for adapting language models to new safety tasks from a compact task specification consisting only of a natural-language description and a handful of positive and negative examples. We formulate task-specific safety adaptation as complementary read and write operations over a shared hidden-representation space: task-relevant risk is read from frozen representations for prediction, while task-specific behavioral changes are written through activation interventions. Starting from minimal task-specific supervision, UniTrust trains a lightweight detector over frozen hidden representations to identify task-specific risks and derives activation directions from matched pairs of safe and unsafe responses to guide generation. To provide broader task coverage without additional human annotation, our Corpus-Grounded Synthesis procedure retrieves diverse contexts from public corpora and rewrites them according to the target task specification, constructing supervision for both detection and intervention. With only ten positive and ten negative real examples per task, UniTrust consistently improves detection across ten safety domains and three model families, and reduces average unsafe compliance by more than 50% across aligned models and their safety-weakened counterparts. Experiments demonstrate that effective safety adaptation can be built from a minimal task specification with little task-specific human supervision. Code and data are available at https://anonymous.4open.science/r/UniTrust.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.