acceptodds
Under review as a conference paper at ICLR 2027

StateWise: Teaching Language Models to Track and Obey Evolving Constraints in Multi-Turn Dialogue

Abstract

Users of conversational assistants rarely state all requirements up front: they add constraints, amend them, withdraw them, digress, and return. Instruction-tuned large language models (LLMs) degrade in this regime in two ways—active constraints stop being honored as turns accumulate (persistence failure), and withdrawn constraints keep being enacted (relapse)—and the only existing repair we are aware of is an external ledger that re-injects a compiled specification at every turn. To tackle this, we propose StateWise, a self-improving approach that teaches a single LLM to maintain its own revisable constraint state over a typed state with tombstones for revoked constraints, so that persistence and relapse are measured by executable checkers and ground-truth states are compiled for free. StateWise first fine-tunes the LLM on two complementary skills, summarizing a dialogue into its constraint state and responding conditioned on that state; it then self-trains on trajectories the model itself generates, keeping only turns that pass an exact round-trip check against the compiled state. This recipe teaches 7–8B models to track the state almost perfectly, but, surprisingly, not to obey it: on Llama-3.1-8B every supervised variant roughly doubles relapse, and on Qwen2.5-7B none reduces it. We therefore add a preference stage on self-generated relapse contrasts, in which the rejected response enacts a revoked rule and the chosen response passes the round-trip check. Experiments on two backbones with held-out constraint families and unseen phrasings show that StateWise cuts relapse by 8–22 points relative to supervised self-training, lifts conversation success by 8–19 points over the backbone, and beats an oracle compiled-ledger baseline on relapse with no inference-time scaffolding. We also characterize the recipe’s cost on cumulative-constraint and single-turn benchmarks and study weight interpolation as a mitigation.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.