CAIC: A Decision-aware Intervention Framework for Counterfactual Sensitivity in LLMs
Abstract
Large language models can change consequential decisions after demographic- only edits to otherwise identical inputs. We introduce Counterfactual Activation-based Intervention with Certification (CAIC), a white-box frame- work for reducing decision-level counterfactual sensitivity and bounding what remains. Through a Locate–Forecast–Select–Verify pipeline, CAIC admits low- rank activation interventions only when they satisfy task-specific requirements at the deployed scoring endpoint, and abstains otherwise. We instantiate two op- erators: CAIC-S, a spectral intervention that removes dominant counterfactual directions, and CAIC-T, a task-protected intervention that limits suppression of task-relevant signal. On a controlled task, CAIC-S reduces the mean counterfac- tual gap by 96.6%, with a 95% population upper confidence bound of 0.030. On a semi-synthetic salary-screening task with entangled counterfactual and task sig- nals, CAIC-S reduces final-output raw gaps by 69%–78% across two 7B models, attaining the lowest residual gap among seven matched baselines, though partly through score compression. As an early-exit decision rule at its intervention site, admitted CAIC-T reduces the gap by 32%–74% while better preserving task signal and score structure; however, the mitigation does not persist through downstream computation and CAIC abstains there. Our results underscore that counterfactual intervention should be selected and evaluated at its deployed scoring endpoint.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.