acceptodds
Under review as a conference paper at ICLR 2027

Aligning Judge Models to Be Logically Consistent Multi-rule Interpreters

Abstract

Technical governance of AI models will soon require training them to follow thousands of rules. Even making agents follow the law requires them to reason about legal codes from hundreds of countries. However, as we show in this work, presenting an entire ruleset to one judge model—rather than having separate judges interpret one rule each—changes the way many models interpret the requirements. When evaluating the ruleset as a whole, a judge model can reach a conclusion that conflicts with its rule-by-rule judgments. Such inconsistencies could cause misalignment by leading a model to treat a hard rule as a soft constraint. For example, we find that when a rule banning pursuing non-specified objectives appears in a larger ruleset, judge models are less likely to deem it violated. As in human statutory interpretation, surrounding rules shape models' interpretations. We argue, however, that these multi-rule interactions must be measured and handled explicitly. We examine several methods for doing so, including prompting, iterative rule refinement, and consistency training. We hope this work helps establish standards for multi-rule alignment and prevents undetected shifts in interpretation from causing downstream alignment failures.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.