EZCAM: Data, Infrastructure, and Policy for Embodied Photography
Abstract
Embodied photography asks an agent to move a camera, observe the resulting view, and retain a photograph under a limited interaction budget. Studying this capability requires both action-linked supervision and an environment that exposes the consequences of the agent's own decisions. We present EZCAM, connecting photographic trajectories, interactive camera infrastructure, and a learned reference policy. The corpus contains 96,700 video–pose trajectories from games and game-derived reconstructions, with 347,111 auxiliary Internet clips indexed separately. Policy training also incorporates Internet video. Replay audits distinguish improving endpoints from non-monotonic local motion, and a blinded study supports selected data-image pairs. The infrastructure provides 49 scene assets with explicit camera conventions, complete-path checks, and observation accounting. We evaluate seven trained configurations across all 49 scenes, using six starts per scene and 40 action requests per episode. The learned policy achieves the highest mean best-observed raw score gain of 0.5126, compared with 0.2817 for the strongest supplied visual-feature adapter. Its mean endpoint gain is negative, illustrating why retained-image quality and final-pose quality must be distinguished. These results establish a measurable data–interaction–policy pipeline and characterize the trained-model comparison. The evaluation reports training, development, and unseen scenes separately. Controlled ablations measure the effects of branch supervision and recovery; independent policy-output preference remains unverified.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.