acceptodds
Under review as a conference paper at ICLR 2027

CamoWorld: Towards Image World Modeling for Camouflaged Object Understanding

Abstract

Camouflage is a natural concealment mechanism where objects blend into their surroundings and become difficult to distinguish from complex backgrounds. Humans rely on internal world models to infer unobserved states, understand scene regularities, and form plausible expectations from limited observations. In this paper, we present CamoWorld, the first self-supervised latent image world model for camouflaged object understanding. Rather than fitting task-specific outputs or reconstructing pixels, CamoWorld formulates camouflage understanding as latent predictive world modeling. Specifically, it introduces three latent world variables: 1) a mask variable for local structure completion, capturing fine-grained textures, boundaries, and subtle structural discrepancies; 2) a spatial relation variable for global scene continuation, modeling spatial organization and continuity between hidden objects and surrounding backgrounds; and 3) a domain variable for appearance prediction, capturing predictable changes across imaging conditions. CamoWorld learns camouflaged world knowledge from raw images within a unified latent prediction framework. Qualitative and quantitative analyses show its ability to capture such knowledge, while extensive downstream experiments demonstrate that latent world prediction learns more effective representations than existing visual pre-training paradigms.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.