Lens as Action: World-State Querying for Scene-Consistent Photographic Generation
Abstract
Camera-controllable text-to-image generation aims to synthesize photographs that comply with both semantic prompts and specified camera settings while maintaining scene coherence. Existing methods typically inject lens parameters directly into the denoising process as auxiliary conditions, but their responses to continuously varying parameter values can be weak or non-monotonic. Because scene content and camera-dependent appearance are represented jointly, changing a lens setting may also inadvertently alter scene geometry or semantic layout. In this work, we revisit photographic image generation from the perspective of lens-action world modeling. Rather than treating lens parameters merely as generative conditions, we interpret them as observation actions over a latent scene world. Based on this formulation, we propose a Lens-Action World Model (LAWM) for scene-consistent photographic generation. Given a text prompt, LAWM establishes an anchor observation and encodes its appearance together with geometric cues into a persistent world state shared across lens queries. A structured Lens-action encoder and World-Action Cross-Attention retrieve action-relevant features from this state and inject them into a temporally coupled diffusion backbone, producing photographs under the requested settings. To alleviate the scarcity of aligned multi-setting photographs, we construct training supervision for lens parameters such as focal length and aperture using field-of-view transformations and depth-guided defocus simulation. Extensive experiments show that our method achieves stronger scene consistency and more reliable responses to continuous lens controls, demonstrating the potential of explicit world-state querying for camera-controllable image generation.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.