acceptodds
Under review as a conference paper at ICLR 2027

EZCAM: Data, Infrastructure, and Policy for Embodied Photography

Abstract

Embodied photography asks an agent to move a camera, observe the resulting view, and retain a photograph under a limited interaction budget. Studying this capability requires both action-linked supervision and an environment that exposes the consequences of the agent's own decisions. We present EZCAM, connecting photographic trajectories, interactive camera infrastructure, and a learned reference policy. The corpus contains 96,700 video–pose trajectories from games and game-derived reconstructions, with 347,111 auxiliary Internet clips indexed separately. Policy training also incorporates Internet video. Replay audits distinguish improving endpoints from non-monotonic local motion, and a blinded study supports selected data-image pairs. The infrastructure provides 49 scene assets with explicit camera conventions, complete-path checks, and observation accounting. We evaluate seven trained configurations across all 49 scenes, using six starts per scene and 40 action requests per episode. The learned policy achieves the highest mean best-observed raw score gain of 0.5126, compared with 0.2817 for the strongest supplied visual-feature adapter. Its mean endpoint gain is negative, illustrating why retained-image quality and final-pose quality must be distinguished. These results establish a measurable data–interaction–policy pipeline and characterize the trained-model comparison. The evaluation reports training, development, and unseen scenes separately. Controlled ablations measure the effects of branch supervision and recovery; independent policy-output preference remains unverified.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.