acceptodds
Under review as a conference paper at ICLR 2027

FREEZER: Extracting More from Frozen Representations

Abstract

A pretrained encoder may retain useful task information that its scoring rule does not fully use for the population being evaluated. We introduce FREEZER (Feature Reweighting and Embedding Enhancement with Zero Encoder Retraining), a post-hoc framework that uses dataset-level embedding statistics and supervision, when available, to adapt scoring while keeping the encoder fixed. For speaker verification, it retains cosine similarity after centering and covariance-based reweighting. For speech and image deepfake detection, it fits a regularized scoring direction and optionally blends the original detector score. On Common Voice 26, macro equal error rate (EER) across 51 locales falls from 3.03% to 2.17%, with improvements in 49 locales. A separate collective face-verification procedure reduces ArcFace IJB-C EER from 1.20% to 0.95%, using unlabeled evaluation embeddings and neighboring comparisons. Speech deepfake detection EER falls from 20.79% to 10.05% on the full CoSG release, with adaptation using only separate MLAAD and In-the-Wild data. Image deepfake detection with AIDE reduces EER from 10.29% to 7.53% on the CNN benchmark. Component studies show that supervised centering supplies most of the speaker-verification gain, while a fitted direction preserves most of the speech deepfake detection gain without blending. For image deepfake detection, blending helps some populations but hurts others. These findings demonstrate the value of adapting the scoring stage of frozen systems, while showing why gains must be evaluated across populations and operating points.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.