PanoReward: Adaptive Multi-Criteria Reward Modeling for Panoramic Image Generation
Abstract
Panoramic image generation has become an important component of VR/AR, cinematography, and world modeling. However, progress in this area is constrained by the lack of a reliable reward model, which in turn caps achievable performance. We posit that reward modeling for panoramic generation should be inherently multi-criteria, explicitly integrating multiple evaluation dimensions and adopting adaptive, preference-conditioned weighting to reflect judgment standards across diverse task requirements. To mitigate this bottleneck, we propose PanoReward, a reward model that evaluates panoramic output along three complementary axes: (i) 3D consistency, (ii) polar visual quality, and (iii) image-text consistency. The model is trained on PanoReward-Data, which contains 943,074 task-profile training instances derived from 195,930 source pairs. In panoramic image generation settings, PanoReward exhibits a strong and consistent alignment with human preferences under different task requirements. In addition, PanoReward supports downstream workflows for panoramic generation, e.g., reward-driven post-training. We will release both PanoReward and the associated training dataset.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.