acceptodds
Under review as a conference paper at ICLR 2027

Source Gains, Target Losses: Auditing Geometry-Distilled Visual Features

Abstract

Geometric supervision promises visual representations that better capture scene structure, yet gains on source geometry tasks can hide a loss of boundary performance on other datasets. This creates a selection problem: an adapter can score better on the task it learns while becoming less useful for the task it is meant to support. We study this mismatch by adapting DINOv2 with depth and multi-view supervision from ScanNet++ and fitting new linear boundary probes on three labeled target datasets. With adapter capacity, training objective and budget held fixed, adaptation depth separates the outcomes: early-layer updates transfer better than late-layer updates in all three training runs, under five evaluation metrics and two dataset weightings. Source boundary F1 nevertheless favors a later layer group in every run, and the chosen adapters score below the unadapted backbone. We then test whether geometric learning can keep more of the backbone's boundary performance by penalizing changes to pretrained patch features. With a penalty weight chosen using target results, the mean target F1 loss falls from 0.015–0.020 in three unconstrained training runs to 0.0010 on average over three constrained runs, while source coplanarity gains remain. A preregistered test on COCO, a target set the project had never used, repeats the layer result: early windows beat both late windows and the unadapted backbone in all three runs, and source boundary F1 again picks a worse window. Adapting a backbone well therefore depends on where it learns geometry and on which pretrained features it keeps for transfer.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.