VLMs Control the View: Learning to Plan Camera Movements through Self-Exploration
Abstract
Can VLMs predict how camera movements change the view, and learn to choose those movements to find a target view? We study this capability as view planning, separating tracking a supplied path from composing one through interaction. We introduce ViewSuite, a benchmark with three tasks that pair supplied-path understanding and interactive camera control in real 3D scenes. Across 13 VLMs, the strongest models exceed 70% on short-distance tracking tasks but reach at most 21.3% on interactive view planning. Learning Interactive View Planning is difficult when the base policy rarely succeeds: Qwen2.5-VL-7B begins at 2.5% success, and direct PPO and GRPO reach only 3.2% and 4.9% under the tested recipes. Our key observation is that every exploration trajectory, successful or not, records valid transitions between views. View Graph Distillation accumulates these transitions across episodes and connects them at nearby camera poses. It samples paths whose final views become new targets, while the recorded actions and intermediate observations provide demonstrations for those targets. The same graph supplies auxiliary view-understanding supervision. Alternating this distillation with reinforcement learning improves Qwen2.5-VL-7B's IVP success from 2.5% to 47.7% on held-out scenes. The trained checkpoint is also a better starting point for related view-understanding tasks.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.