acceptodds
Under review as a conference paper at ICLR 2027

CARAE: Matched Counterfactual Transition Learning

Abstract

Counterfactual supervision can make a model respond to an edited input, but the response may reflect nuisance variation introduced by the edit rather than the intended task change. A targeted counterfactual alone provides no reference for this ambiguity. Matched controls turn absolute response fitting into a relative comparison of target and non-target transitions within the same input. CARAE instantiates this idea with a question-conditioned visual state, a relative transition objective, and question-conditioned token-weight-modulated residual coupling to a vision–language model’s generation pathway. Expert localization in chest radiographs permits controlled target and non-target edits; the masks construct interventions offline and are never model inputs. On the canonical CheXlocalize protocol, CARAE raises target-versus-control answer AUROC from 0.5133 to 0.8038 at 3B scale and from 0.7053 to 0.8527 at 7B scale. The 7B paired improvement is 0.1474 (95% patient-bootstrap interval [0.0737, 0.2266]). Frozen checkpoints retain an advantage under held-out control constructions and replacement operators; dataset transfer within chest radiography is moderate on MS-CXR and weak in absolute terms on VinDr-CXR. These results support intervention-specific response under the evaluated construction family, without treating synthetic removal as a clinically adjudicated diagnosis.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.