acceptodds
Under review as a conference paper at ICLR 2027

Semantic Inductive Bias for Disentanglement through Vision-Language Models and Iterative Refinement

Abstract

Disentangled representation learning aims to recover meaningful factors of variation from observations. However, the observations alone do not uniquely determine a single factorization, making inductive bias necessary. The factors typically used in disentanglement, such as shape, color, pose, and position, are semantic concepts selected by humans. Recovering these factors benefits from inductive bias informed by semantic knowledge. Pretrained vision-language models (VLMs) acquire broad visual-semantic knowledge through large-scale vision-language pretraining and can provide this form of inductive bias. We therefore formulate factor discovery as a set-level comparative task, presenting small collections of unlabeled images as contact sheets from which a VLM identifies candidate semantic factors and provides visual evidence for each one. Rather than accepting these proposals directly, we iteratively validate and refine them according to sufficiency, minimality, and semantic independence using complementary image-based evidence. We evaluate the method on four disentanglement benchmark datasets using only 300 unlabeled target images and locally deployed open-weight models. The method consistently recovers the ground-truth factors on XYCS, dSprites, and 3D-Shapes, while MPI3D-C produces more variable semantic decompositions. The resulting representations are competitive with -VAE, DAE, and InfoGAN trained using the same 300 images and remain competitive with baselines trained using all images in the target dataset.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.