acceptodds
Under review as a conference paper at ICLR 2027

CAIC: A Decision-aware Intervention Framework for Counterfactual Sensitivity in LLMs

Abstract

Large language models can change consequential decisions after demographic- only edits to otherwise identical inputs. We introduce Counterfactual Activation-based Intervention with Certification (CAIC), a white-box frame- work for reducing decision-level counterfactual sensitivity and bounding what remains. Through a Locate–Forecast–Select–Verify pipeline, CAIC admits low- rank activation interventions only when they satisfy task-specific requirements at the deployed scoring endpoint, and abstains otherwise. We instantiate two op- erators: CAIC-S, a spectral intervention that removes dominant counterfactual directions, and CAIC-T, a task-protected intervention that limits suppression of task-relevant signal. On a controlled task, CAIC-S reduces the mean counterfac- tual gap by 96.6%, with a 95% population upper confidence bound of 0.030. On a semi-synthetic salary-screening task with entangled counterfactual and task sig- nals, CAIC-S reduces final-output raw gaps by 69%–78% across two 7B models, attaining the lowest residual gap among seven matched baselines, though partly through score compression. As an early-exit decision rule at its intervention site, admitted CAIC-T reduces the gap by 32%–74% while better preserving task signal and score structure; however, the mitigation does not persist through downstream computation and CAIC abstains there. Our results underscore that counterfactual intervention should be selected and evaluated at its deployed scoring endpoint.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.