Does My LLM Obey Me? Falsification-Based Attribution Tests for Prompt-Based Behavioral Conditioning
Abstract
Validating the effectiveness of a behavioral specification just by tracking the change in the behavior of a large language model (LLM) relative to a baseline can be deceptive, as the changes may arise from components of the specification other than the intended ones. To address this issue we propose four falsification-based attribution tests - misassignment, reversal, content-irrelevant control and attribute name substitution. We apply these to a behavioral policy system for emotional support dialogue and show that it passes its baseline test decisively. The fraction of response sentences that are considered advice goes down from when the policy is added with a low value of the "advise" parameter . On Qwen2.5-3B, with partial replications on Phi-3.5-mini and Mistral-7B, the central comparison survives independent human annotation of 40 conversations with a bias-corrected reanalysis. Renaming eight attributes without changing their values raises advice proportion by relative to the original names. Inverting all eight values (turning low values to high and vice versa) leaves it indeterminate under human annotation as under a classifier . Their within-conversation difference is . The conclusion remains unchanged under blind adjudication. The pattern is not specific to advice, and the same ordering appears on reflection. Relative sensitivity to the tested interventions is configuration-specific. On a non-counseling counted-outcome task, changing the value of "verbosity" from increases response length . Verbal encoding of the values as qualifier words instead of values (advise: becomes advise: rarely) partially rescues value sensitivity on advice . We investigate the relative name/value sensitivity changes across Qwen2.5 model configurations and prompt placements. We also find evidence against assuming that a specification validated in one configuration remains valid in another, and release the tests as "speccheck".
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.