CooperTrace: Cooperative Multi-Modal End-to-End Motion Prediction
Abstract
Modern autonomous driving stacks transition from modular pipelines to end-to- end learning frameworks. Equipped with diverse sensing modalities, autonomous vehicles are increasingly robust to sensor degradation and long-tail events. In the meantime, leveraging Vehicle-to-everything (V2X) communication, recent cooperative autonomy research also expands from perception to downstream tasks, demonstrating the benefits of situational awareness beyond individual sensing range. However, it is unclear how multi-modal sensing would interact with the end-to-end learning paradigm in the V2X setting. In this paper, we present CooperTrace, a cooperative, multi-modal, and end-to-end motion prediction framework to explore the feasibility and efficacy of this three-fold combination. CooperTrace fuses LiDAR and camera at every agent, attends to the features of connected agents, and trains the forecaster on tracking results from previous frames. Extensive evaluations on both simulated and real-world cooperative datasets show that CooperTrace achieves up to 68% lower motion prediction error, 134% higher MOTA, and 18% higher 3D detection AP than the compared state-of-the-art methods. We further stress-test its cooperative gains under emulated sensor degradation. Cooper Trace’s multi-modal multi-agent fusion maintains stable performance across these settings and degrades gracefully when one modality becomes unavailable, demonstrating its potential for harsh real-world deployment conditions
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.