acceptodds
Under review as a conference paper at ICLR 2027

What matters for perceiving object-parts in vision encoders?

Abstract

Recent works provide evidence that large-scale pretraining makes vision and language models organize data similarly, however this similarity is exclusively studied on scene- and object-level. This leaves the unanswered question: are object-parts organized in vision models similar to language models? Here, we focus on parts, as user-defined components of 3D objects. Our work tries to answer this question through a series of analyses: **(i) Part-similarity.** We measure part-level representation similarity between frozen vision and language models using a modality-agnostic protocol. To this end, we conduct the first large-scale comparison of 39 vision encoders across 3 different vision modalities. **(ii) Part-alignment.** We further propose to test the ease of transforming vision models to language in the absence of parallel part annotations. To this end, we propose a simple probe based on differentiable optimal transport, particularly suited for data-scarce settings like 3D where large-scale paired part annotations are not available. **(iii) Part-transferability.** Finally, we find that our probe can transfer to unseen part hierarchies when trained on bags of discovered part masks and cheaply obtained VLM-provided part names. Our analyses demonstrate how visual pretraining strategy, architectural inductive biases, and modality influence both similarity and alignment with language on a part level with implications towards real-world applications.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.