Per-Sample Routing of Pre-trained Encoders via Embedding Typicality
Abstract
Practitioners typically deploy pre-trained vision encoders by choosing a single backbone and layer, fitting a probe on its embeddings, and using that representation for every test sample. This fixed choice can be suboptimal because the best encoder and layer can vary across inputs in heterogeneous tasks. We study unsupervised per-sample routing over a pool of frozen pre-trained encoders, treating each (backbone, layer) pair as a candidate. Using only embedding geometry, the router either assigns a test sample to one candidate or combines predictions across the pool. It requires neither supervised router training, test labels, nor per-candidate calibration. For hard routing, we repurpose the existing Levy radial-typicality score: it reads typicality from the embedding norm and selects one candidate per sample. For soft ensembling, we introduce the local spectral normalized gap (LSNG), which measures local spectral typicality and produces dimension-comparable weights over the full pool. On a pool of 192 candidates from 2 general-purpose backbones and 6 finetuned experts, LSNG soft ensembling gains +7.8 to +54.1 percentage points (pp) in accuracy over the DINOv2-L last-layer default under a linear probe. This is within 1.0 pp of the test-selected best fixed candidate on 3/5 evaluations and above it on the other 2. Levy hard routing evaluates 1 probe per sample, matching the fixed default's probe cost, and improves accuracy by up to +31.1 pp on 3 evaluations. It fails when strongly finetuned off-domain experts imitate typical embedding norms and attract the argmax; soft ensembling is less sensitive to this failure. Across this heterogeneous library, training-free embedding geometry provides enough signal to route between frozen encoders.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.