acceptodds
Under review as a conference paper at ICLR 2027

CC-TTRL: Control-Conditioned Test-Time Selection of Normalization Statistics under Label Shift

Abstract

Test-time adaptation methods based on Batch Normalization can correct batch-specific technical variation using unlabeled target samples, but their statistics become biased when the target class distribution shifts. In biomedical imaging, negative controls provide a stable reference for separating technical variation from perturbation-specific signal. CS-ARM-BN exploits this structure by pooling controls and perturbed samples when recomputing BatchNorm statistics, yielding a favorable global bias-variance trade-off. However, the same estimator is applied to every normalization layer and every target batch, despite the fact that class-dependent activation bias varies across network depth and target composition. We introduce CC-TTRL, a control-conditioned test-time adaptation method that selects the source of normalization statistics independently across network stages. For each stage, CC-TTRL chooses among pooled target statistics, control-only statistics, and stored training statistics, producing a discrete family of adaptations that includes CS-ARM-BN as a special case. Candidate configurations are evaluated without target labels using a cross-fitted reward that combines agreement between held-out target controls and source controls with prediction consistency on held-out perturbed samples. A trust-region constraint and improvement veto preserve the CS-ARM-BN solution when evidence for an alternative configuration is insufficient. We search this action space using a lightweight policy-gradient procedure. We evaluate CC-TTRL on mechanism-of-action prediction from high-content cellular microscopy under experimental-batch shift and label shift. Under severe label shift, CC-TTRL improves over CS-ARM-BN and preferentially avoids pooled statistics in deeper network stages, consistent with increased class-dependent activation leakage at greater depth. The proposed control-anchored reward correlates with held-out predictive accuracy and substantially outperforms a feature-space reward. Random search with the same reward and evaluation budget achieves comparable performance to policy-gradient search, indicating that the primary contribution is the layerwise adaptation space and label-free control-guided selection criterion rather than reinforcement learning itself.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.