Data-Efficient Criteria Steering for LLM Judges
Abstract
LLM judges are increasingly used to enforce safety and compliance policies, but adapting them to new languages, regulations, and risk taxonomies remains costly. Fine-tuning requires substantial labeled data and model updates, while prompt optimization relies on representative evaluation data and can overfit when data are scarce and the signal is limited, reducing held-out performance. We introduce criteria steering, a data-efficient, training-free method for specializing LLM judges from small labeled datasets without updating model weights. Our pipeline mines and ranks candidate criteria, constructs a contrastive activation direction from diverse criteria, and applies it only to the criterion-token span during prompt encoding, adding no decoding-time computation. On Bengali toxicity, Japanese content safety, and Indian financial-services safety and compliance (FinProof), criteria steering improves held-out performance for small general-purpose as well as purpose-built guardrail models. It outperforms text-only prompt optimization and, on FinProof, matches or exceeds a strong automated prompt-optimization reference depending on the operating point. Gains persist with relatively small labeled samples. These results position criteria steering as a practical adaptation mechanism for low-resource languages and specialized domains.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.