acceptodds
Under review as a conference paper at ICLR 2027

CALIBRE: Fold-Calibrated Readouts for Frozen Pathology Encoders

Abstract

Frozen pathology encoders are evaluated on spatial gene expression prediction, but fixed readouts can obscure their value. HEST-Benchmark's ridge penalty leaves 256.0 of 256 effective degrees of freedom. We introduce CALIBRE, which adapts regularisation by selecting each encoder's penalty and projection dimension on training slides, averages seven encoders and applies an existing within-slide smoother. Calibration improves all seven pre-registered backbones on nine tasks. CALIBRE scores 0.4664 and the smoothed uncalibrated average 0.4575, versus 0.4480 from STFlow's published task scores. On 19 held-out breast cancer patients, calibration raises the primary encoder from 0.0811 to 0.1033 (p=0.0039). CALIBRE scores 0.1213, versus 0.0765 for STFlow, a flow-matching model with training-selected epochs (p=0.00064), or 0.0580 with a tuned learning rate (p=0.00042). On a 22-patient kidney cohort from another institution and platform, registered before download, calibration improves the primary encoder for every patient and CALIBRE exceeds STFlow, 0.2144 versus 0.1795 (p=0.000009). Predictors sharing encoder and smoother separate only against STFlow weakened by tuning. HER2ST's 0.2644 versus 0.2622 remains unresolved, and averaging is resolved on one of four cohorts. Training-selected frozen readouts change frozen-versus-trained comparisons.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.