Intervention Effects Are Execution-Relative
Abstract
Representation-level interventions are commonly compared and tuned through downstream task performance. Yet an intervention-induced state does not directly produce a task output: it is converted by a final execution rule, such as thresholding, decoding, suppression, or candidate scoring. We ask whether conclusions about an intervention family remain stable when this downstream rule changes while the intervened states are held fixed. We call this dependence execution-relative intervention effects and introduce matched-state evaluation, which computes each intervention state once and reuses the same cached state across task-valid executors. Across sound, language, vision, and speech tasks, changing only the executor can alter the measured effect size and, in stronger cases, reverse the ordering of intervention settings or change the selected setting. In sound event detection, fixed-threshold and hysteresis execution give opposite local orderings to the same intervention transition and select different settings under cross-fitting; transferring the fixed-threshold choice to hysteresis incurs a Event-F1 loss on held-out data. Boundary manipulations trace the fixed-threshold drop to sparse crossings of event-formation boundaries, revealing candidate instantiation as one mechanism. A complementary Qwen3-8B study keeps candidate statistics fixed and changes only the scoring rule, showing how candidate scoring reshapes competitive margins and intervention responses while preserving the selected coefficient. These results show that intervention evaluation and selection should account for the executor that turns intervened states into task outputs, and should test whether comparative conclusions remain stable across plausible task-valid execution rules.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.