MATRICS: Protocol-Equivalent Attribution of Clinical Text Pipelines for Causal Subgroup Discovery
Abstract
A text representation in causal subgroup discovery helps define the hypotheses that are ultimately tested: changing it can rotate proposal directions, move patients across score tails, and alter the effective search family. MATRICS isolates this consequential intervention by passing each representation pipeline through the same leakage-free operator for cross-fitted doubly robust scores, sparse signed-tail proposals, ranking, Benjamini–Hochberg screening, and complete-path stability selection. On 50 paired semi-synthetic MIMIC-IV replicates, the modifier-adapted ClinicalBERT pipeline recovers known note-borne effect modifiers at F1 0.691 versus 0.585 for TF–IDF, a gain of 0.106 (18.1% relative; paired 95% interval [0.061, 0.148]; 39/50 wins); CATE RMSE falls from 0.301 to 0.243 and matched-predicate stability rises from 0.73 to 0.86. The reference configuration and all six one-axis alternatives retain positive lexicalized intervals. Crossed information regimes explain the gain: masking literal modifier spans preserves a 0.056 advantage, every structured-only contrast lies within a four-point recovery-F1 compatibility band, and global-null runs produce at least one stable survivor in 10/50 versus 8/50 replicates. Together with a representation ladder that places domain-aligned biomedical encoders at the strongest observed operating points, these results show when contextual clinical text changes recoverable effect heterogeneity: it matters most when notes carry modifier information not duplicated in structured covariates.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.