R0: Feed-Forward 3D Reconstruction for Arbitrary Central Cameras
Abstract
Autonomous driving, robotics, and immersive VR/AR systems increasingly rely on fisheye and panoramic cameras for perception. However, existing feed-forward 3D reconstruction models are built on pinhole projection and degrade significantly on non-pinhole cameras. The root cause is that different cameras are treated as discrete categories, whereas all central cameras actually share the same ray space and differ only in how they unfold rays onto the image plane. We present R0, to our knowledge the first camera-agnostic feed-forward reconstruction model trained at scale across pinhole, fisheye, and panoramic cameras. By predicting a continuous ray field, constructing spherical neighborhood sampling, and conditioning cross-view attention on relative geometry, R0 unifies sampling, interaction, and output before projection. It natively handles different fields of view and unfolding functions, without requiring camera intrinsics, projection type, or undistortion at inference. Experiments on pinhole, fisheye, panoramic, and heterogeneous cameras show that R0 achieves state-of-the-art results on non-pinhole data while remaining competitive on pinhole data. Code, models, and large-scale data processing details will be released.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.