Programmable Semantic Interventions: Learning Causal Carriers from Behavioral Programs
Abstract
Causal abstraction explains a frozen network by aligning its activations with the variables of a simpler, interpretable model, and checking that interventions on the two agree. Finding such an alignment currently requires a counterfactual dataset: for every pair of prompts, the answer the network ought to give once the variable has been changed. Producing those labels means committing to an executable algorithm for the behavior before studying the model supposed to implement it. However, the criterion these methods optimize only leverages the answer an intervention should produce. Here, we developed an approach that lets the investigator define that relationship directly: which answers the variable is about, and how they should move when it changes. We call this interface a *semantic program* and introduce *Programmable Semantic Interventions* (PSI, ): a low-rank activation subspace learned so that interchanging it moves the model's answer the way the program declares. On the MIB benchmark matches or surpasses the quality of supervised alignment, while requiring minimal supervision. The declaration also fixes a coordinate system with known semantics, enabling donor-free steering: jointly training for both localization and control realizes % of held-out requests for answers the investigator chooses, without the need for donor prompts.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.