Behavioral Spec Inversion: Recovering a Model's Enacted Specification from a Limited Record of Its Behavior
Abstract
A model specification describes how an assistant should behave, but the specification a model enacts, request by request, can differ from the one written for it and from the one it reports when asked. The enacted specification is what auditors and developers need, yet an adherence score and a list of failures report how often a model fails, not when. We study behavioral spec inversion (BSI): recovering a model's enacted specification from a limited record of its black-box behavior, as a natural-language description that is selected by behavioral evidence and tested on behavior it was not produced from. On value-adapted targets whose intended preference is known, and on two targets trained on specifications we wrote, a recovered description reaches the preference more reliably than the specification's parameters: only some supplied preferences are identified stably, and a description written from a small set of choices misstates a threshold yet reproduces most of the target's choices on unseen probes. On SpecBench Code with natural targets, behavior profiles learned by BSI-Search from a small number of observations predict clause violations on test queries, with every prediction frozen before the responses exist: they improve on clause frequencies and on direct rule induction, mostly by anticipating which clauses a request engages, and an implanted condition is recovered and attached to its correct branch. Rendered into corrective training data at equal budget, a profile raises specification adherence over frequency-guided data on a natural target, at a cost in benign answering, and repairs an implanted defect no better than ordinary compliant examples.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.