IDOL: Answer-Free Behavioral Constraints for Small Language Models
Abstract
Behavioral constraints can guide a small language model without supplying a target response, but their contribution can be conflated with changes in prompting, candidate generation, and evaluation. We introduce IDOL, a source-labeled framework for measuring the incremental value of answer-free behavioral guidance. Typed IDOL-DSL programs distinguish requirements recoverable from the prompt from additional behavioral proposals, and compile into deterministic checks and semantic verifiers over student-generated responses. We evaluate canonical and prompt-derived programs through separate generation interventions and fixed-pool selector comparisons, using two automated judges in both response orders. On a fixed, preselected 300-prompt UltraChat-derived panel across three generation seeds, canonical guidance improves automated scores over a generic sham under both judges, while the closest explicit and permuted controls do not establish an advantage. Accepted latent constraints occur in 66 prompts and change 28 of 900 shared-pool selections relative to prompt-derived scoring. The targeted dual-order selector estimate is under DeepSeek, with a 95% interval of , and under Phi, with : both point estimates are positive, but the declared cross-judge support criterion is unmet. Controlled, filtered-rationale, and benchmark extensions yield inconclusive or mixed results. IDOL provides an executable representation and an attribution-focused study of constraint-guided inference, distinguishing a prompting gain from the incremental effect of additional scoring constraints.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.