KOMPAS: Koopman Dynamics Modeling and Black-Box Predictive Control of Multi-Turn LLM Reply Attributes
Abstract
In tasks such as multi-turn jailbreak defense or format-constrained generation, a language model must hold a reply attribute at a target on every turn. However, the attribute does not reset between turns: what one reply does on it carries into the next, so its readings form a trajectory with inertia. Existing remedies react to the current reading alone and ignore this inertia, so their corrections arrive a turn late or push past the target. We fill this gap with KOMPAS, KOopman dynaMics modeling and Predictive control of Attributes at the prompt Step, which treats a prompt–response exchange as one step of a controlled system. It fits a linear dynamics model to past trajectories of the attribute, with the instruction as the input. Before sending an instruction, KOMPAS predicts how each candidate would move the attribute and sends the one predicted to reach the target, without querying the language model. On three public instruction-following benchmarks, KOMPAS raises constraint satisfaction by 14.5 and 14.8 points over two uncontrolled models. It also beats the stronger of two baselines that steer the model’s internal activations, by 5.3 and 5.7 points.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.