acceptodds
Under review as a conference paper at ICLR 2027

Constitutional Value Potentials: Reading Priority Margins and Steering Value Trade-offs in Language-Model Activations

Abstract

Model constitutions specify how a language model should trade off values such as helpfulness, harmlessness, and honesty when they conflict. We ask whether different value conflicts share internal structure. We introduce Constitutional Value Potentials (CVP), which assigns each value a loading on a shared activation basis and reads a conflict through the difference of two loadings. Jointly learned loadings compose readouts for pairs without their labels. Across five instruction-tuned models from three families, rank-one composed readouts achieve mean within-orientation AUROC of -. They approach readouts fitted to each target pair on conflicts with helpfulness or autonomy and are weaker on protective-value pairs. With target-pair supervision, the shared axis also supports early monitoring: prompt-state pooled AUROC is – versus – for a text classifier, with transfer to requests without an explicit priority clause. Steering along composed directions shifts judged trade-offs in the intended direction on all eight evaluated pairs in Qwen2.5-7B.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.