acceptodds
Under review as a conference paper at ICLR 2027

BEYOND ARGMAX: A MECHANISTIC STUDY OF SEMANTIC RETENTION IN FROZEN FOUNDATION-MODEL COMPOSITION FOR GENERALIZED FEW-SHOT 3D SEGMENTATION

Abstract

Classical classifier-combination work distinguishes score-level fusion from hard decision voting. We revisit this distinction in a modern regime where independently pretrained, frozen foundation models are composed at inference time for generalized few-shot 3D segmentation. Our contribution is not a new sum rule; we ask a mechanistic question: *how much useful semantic information is lost when heterogeneous sources are collapsed to one class before interaction?* We construct a same-input top- retention intervention that freezes model weights, point clouds, masks, geometry, class vocabularies, scene lists, and fusion rules while varying only the semantic support retained before fusion. On 156 held-out ScanNet200 scenes, full retention reaches harmonic-mean (HM) IoU versus for the frozen top-1 control (, 95% paired scene-bootstrap CI ). The effect independently replicates on all 50 ScanNet++ validation scenes ( vs. HM; , CI ). Most of the gain appears immediately after top-1, while moderate top- can slightly outperform the unrestricted tail. The conclusion survives a GroundingDINO–SAM2.1 sparse-source replacement, alternative fusion operators, tie-policy changes, and source-weight variation. Calibration diagnostics show severe but opposite raw miscalibration of the dense and sparse sources, yet calibration alone does not eliminate the retention advantage. We therefore identify *premature semantic collapse* as a repeatable information bottleneck in heterogeneous frozen-model composition.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.