UniGenStereo: Unified Stereo Video Generation with Flexible Conditioning
Abstract
High-fidelity stereo video generation is crucial for immersive extended reality (XR), robotics, and embodied AI. In practice, the geometric information available at inference time varies widely: dense per-frame depth or disparity is sometimes available, sometimes only the camera baseline is known, and sometimes a monocular clip arrives with neither. Existing methods each commit to one of these cases, and those taking no geometric input cannot control the stereo geometry they produce. We propose UniGenStereo, which synthesizes the right view from a monocular video with one set of weights for all three. Disparity is the camera baseline times the ratio between focal length and metric depth. The ratio does not depend on the baseline, so one prediction serves them all; we predict it from the input video with a lightweight head on the generation backbone. (i) Supplied disparity is used as given and the head is never invoked; (ii) a calibrated baseline scales the predicted ratio; and (iii) with no baseline, a standard interpupillary value stands in. Whatever its source, the disparity warps the input view into the same conditioning channels, so one backbone covers all three and varying the baseline alone controls the synthesized disparity. Experiments show state-of-the-art results in all three settings, including on real-world capture.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.