Bridge Transformer: Beyond Ego-Centered Fusion in Collaborative Semantic Occupancy Prediction
Abstract
Collaborative 3D semantic occupancy prediction enables multiple agents to exchange complementary information, yet, regardless of how many agents participate, the prediction remains a single semantic occupancy grid in the ego frame. The task is thus Many-to-One (M2O) in its output alone. Existing methods impose M2O on fusion as well. The resulting M2O fusion is centered on the ego representation and treats collaborator representations as inputs to be aggregated into it. Each collaborator representation, however, carries the errors of its single viewpoint. We empirically observe that M2O fusion passes these errors into the ego grid uncorrected, so the error rate rises in regions observed only by collaborators. This motivates Many-to-Many (M2M) fusion, which jointly refines the representations of all agents to correct these errors and defers the reduction to the ego grid until the output. To this end, we propose the Bridge Transformer, a feed-forward architecture for M2M fusion. Each agent represents its scene as 3D semantic Gaussians. These Gaussians are grouped, and each group is embedded into hierarchical tokens comprising an anchor token and its member tokens. Inter-agent attention over the anchor tokens of all agents alternates with intra-agent attention over each anchor and its members. Because every agent's anchor tokens serve as queries, this alternation jointly refines all agents' representations. A dual-branch Gaussian decoder then refines and extends each agent's Gaussians for splatting onto the ego grid. Experiments show that the Bridge Transformer outperforms the strongest prior method by 3.88 mIoU and 2.01 IoU. An ablation further shows that M2M fusion corrects more collaborator errors than M2O fusion.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.