Co-Registered World-State Sequences for Collaborative Perception and Prediction
Abstract
Collaborative perception may rely on one fused snapshot to expand each agent’s view, whereas prediction benefits from a sequence of shared world states to model how the world evolves. Each state integrates multi-agent observations from one timestamp in a common frame. Existing methods fall into two groups, neither explicitly constructing such a sequence. The first aggregates time before spatial alignment: each agent encodes its observation history and transmits the representation, which is aligned and fused across agents at the current timestamp. The second compensates for delay by extrapolating delayed messages to the current timestamp, again yielding a single fused snapshot. Both reduce historical information to a current-time representation without explicitly constructing shared states at past timestamps, and their temporal features can conflate object motion with observer motion without a common spatial reference across timestamps. A diagnostic comparison highlights room for improvement: constant-velocity extrapolation from ground-truth historical tracks matches published SOTA. This motivates co-registering and fusing multi-agent observations into a shared world state at each timestamp before temporal aggregation, so that scene dynamics can be modeled in a consistent spatial reference. Following this principle, we introduce CoReST (**Co**llaborative **Re**gistered **S**tates over **T**ime), which aligns observations from each past timestamp to the ego vehicle’s coordinate frame at prediction time and fuses them into a shared world state before temporal aggregation. To support this pipeline, World-State Attention constructs each shared state through visibility-aware masking, source encoding, and pose-tolerant attention, while World-State Difference extracts motion cues from differences between consecutive states in the shared coordinate frame. On V2XPnP-Seq, CoReST achieves EPA scores of 52.8 and 52.2 in the vehicle-centric and V2V settings, respectively, surpassing the published SOTA scores of 48.2 and 40.6, while adding less than 2% to the baseline model’s parameter count.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.