From Behavior to Rules: Discovering Executable Constitutions for Language Models
Abstract
When model providers do not disclose the operational specifications governing refusals, auditors must infer their scope from black-box responses. We study behavioral constitution discovery: recovering a compact natural-language program that a separate surrogate model can execute to reproduce an unchanged source's refusal behavior. We measure fidelity through reconstruction error, the fraction of held-out prompts on which source and surrogate answer-or-refuse decisions differ. Our method actively searches for informative prompts and uses all four joint behavioral outcomes to induce rules. Source refusals supply positive evidence, including boundaries already shared by the surrogate, while source answers constrain each rule's allowed scope. Each candidate is tested with and without inclusion in the complete constitution on boundary and stress probes. Acceptance requires local improvement or non-worsening behavior supported by joint refusals, together with preservation of protected examples and non-increasing validation error. This permits explicit recovery of constraints that produce no immediate disagreement, while testing whether separately induced rules remain valid when composed. Experiments across within- and cross-family model pairs indicate that compact constitutions can achieve low reconstruction error and encode refusal boundaries omitted by disagreement-only discovery. When applied to a fixed surrogate, constitutions discovered for different sources induce source-specific refusal boundaries, demonstrating that the recovered rules capture distinct source behaviors. These findings support executable constitutions as concise, inspectable approximations of provider behavior whose fidelity can be tested without access to the original policy text.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.