JFER: Joint Future Exploration and Reasoning for Interactive Motion Forecasting
Abstract
Interactive motion forecasting provides downstream planning with a small set of joint futures, although the number of possible agent-intention combinations can far exceed the prediction budget. Similar hypotheses may occupy several output slots, leaving other plausible interactions unrepresented. We introduce JFER, a Joint Future Exploration and Reasoning framework that addresses this limitation through structured hypothesis construction and complementary exploration. JFER organizes agent-level intentions into a joint-intention bank and assigns six prediction slots to selected pairs before full trajectory decoding. It then identifies additional intentions using their predicted support and distance from the decoded joint endpoints. Each selected intention guides the adaptation of an existing trajectory pair, producing a complementary hypothesis without another full decoding pass. Temporal and cross-candidate attention refine the expanded set, and reliability-constrained selection retains six joint predictions. On the Waymo Open Motion Dataset (WOMD) Interaction Prediction benchmark, JFER achieves a test Soft-mAP of 0.2868. Candidate-set analyses show that complementary exploration increases joint-future coverage and that part of this gain is retained under the original six-mode output budget, with larger gains at longer prediction horizons.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.