Broader Tasks Prefer Ego-Frame Invariant Representations
Abstract
Computer vision has traditionally advanced through models optimized for individual tasks, such as object detection and semantic segmentation, while recent work increasingly seeks general representations that can support diverse downstream objectives. This raises the question of whether general representations can be guided by structural principles that hold independently of any particular downstream task. We study this question in multi-camera perception through the choice of ego coordinate frame, whose origin and axes are arbitrary. Changing this frame alters the numerical camera geometry while leaving the physical cameras, their relationships, and their visual observations unchanged. For image-projected representations, we show that consistent transformations of the underlying 3D quantities and camera geometry cancel under projection, motivating ego-frame invariance (EFI). We hypothesize that EFI is not universally advantageous for every specialized task, but becomes increasingly beneficial as a representation is required to support a broader range of predictions. To test this hypothesis in a setting where the output itself is not invariant, we study 3D object detection, an ego-frame equivariant (EFE) task whose predicted positions and orientations transform with the ego frame. We randomize the ego frame during training, and progressively broaden the task by increasing the number of target classes. Across nuScenes and Argoverse 2, using three camera-conditioning schemes, we evaluate this effect by held-out test performance obtained from selected checkpoints after training. The evaluation provides evidence that the advantage of ego-frame-invariant relative conditioning tends to grow as the prediction task becomes broader. These results support EFI as an inductive bias whose value grows as representations are required to support broader predictions, rather than as a universal advantage for every task.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.