acceptodds
Under review as a conference paper at ICLR 2027

Post-Extraction Compression of Pathology Foundation-Model Embeddings: Storage, Fidelity, and Downstream Utility

Abstract

Whole-slide learning and retrieval often reuse large stores of patch embeddings extracted by a frozen pathology foundation model. We evaluate compression of 1,294,852 CONCH v1.5 embeddings from 1,701 TCGA glioma slides (877 patients) at magnification. We measure stored bytes, CPU decode time, reconstruction and retrieval distortion, attention stability, and three-class multiple-instance learning (MIL) under patient-disjoint evaluation. Per-channel INT4 achieves measured compression and a five-seed mean macro-AUROC of , compared with for FP32. The patient-clustered 95% interval for the difference between five-seed ensembles is , above the specified engineering margin. Adding one packed residual-sign bit per coordinate lowers mean attention total variation from to and increases top-5% attention-set Jaccard overlap from to , while reducing compression to . Normalization followed by Haar rotation reduces the largest recorded coordinate by , but does not reduce maximum empirical excess kurtosis; the measured per-channel rotated decoder takes seconds per bag, compared with ms for INT4. We derive a bag-size-independent bound on gated-attention perturbation. These results support low-bit storage for this encoder, cohort, and MIL model while showing that attention fidelity and decode cost require separate evaluation.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.