CamWAM: A World-Action Model for Autonomous Photography
Abstract
Autonomous photography aims to enable embodied agents to capture visually appealing photographs without human camera operation, yet remains challenging due to the subjective nature of photographic objectives. Existing approaches typically determine camera actions without explicitly modeling their resulting visual outcomes, limiting their ability to reason about how camera motion affects photographic quality. In this work, we propose **CamWAM**, a world-action model for autonomous photography that jointly models future visual evolution and continuous executable camera motions. CamWAM cascades a video expert with an action expert through a lightweight adapter, allowing predicted visual evolution to provide informative representations for camera motion prediction. To effectively couple the two experts, we develop a three-stage video-action learning strategy that progressively aligns visual evolution with camera motion. We further introduce a scalable data pipeline that reverses videos synthesized from curated high-quality photographs to construct camera trajectories toward desirable views, while using independently constructed 3D Gaussian Splatting environments for cross-domain evaluation. Extensive experiments demonstrate the effectiveness of CamWAM and its strong generalization across diverse scenarios. Our models, training code, and dataset will be publicly available.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.