EpiPE: Epipolar Positional Encoding for Camera-Controlled Video Generation
Abstract
Camera positional encodings provide an efficient way to condition video diffusion transformers on prescribed viewpoints. However, encodings based on homogeneous camera transformations do not directly express the epipolar compatibility of viewing rays, a geometric prior available before the scene content is synthesized. We introduce EpiPE, a relative camera positional encoding that incorporates the epipolar prior into learned attention scores. EpiPE applies the classical Plücker line transformation to learned features, exposing the coupling between translation and viewing direction that underlies epipolar geometry. This places epipolar coupling within a consistent coordinate transformation whose composition structure allows relative camera interactions to factor into per-token operations. The resulting encoding integrates epipolar geometry into standard attention without restricting which tokens may interact. Experiments on PanShot and RealEstate10K show improved camera-pose accuracy with comparable visual quality.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.