Learning Interactive Counterfactual Futures for Risk-Aware Embodied Planning
Abstract
Many existing embodied planning methods learn direct mappings from observations and instructions to actions through imitation of recorded demonstrations. Such mappings do not explicitly account for surrounding agents' responses to alternative ego actions, potentially leaving interaction risks unanticipated. We propose a framework that learns interactive counterfactual futures for risk-aware embodied planning. Given multimodal observations and language instructions, a vision-language backbone generates candidate ego trajectories. For each candidate, a learned interaction graph encodes inter-agent relationships, and a meta-action-guided response predictor models surrounding agents' speed and path behaviors before decoding them into future trajectories. The resulting candidate-specific future scenes enable explicit risk assessment and trajectory selection. Training proceeds in two stages: task module pretraining first learns interaction relationships from recorded trajectories; LoRA adaptation then aligns perceptual representations with the planning objective by supervising the risk-aware ego trajectory against the recorded ego future. On Bench2Drive, our method with a Qwen2-0.5B language backbone achieves a Driving Score of 87.47 and a Success Rate of 72.73%, exceeding the strongest compared baseline by 9.43 points and 17.64 percentage points, respectively. The full framework also improves aggregate planning scores over the variant without interaction analysis, with modest reductions in efficiency and comfort.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.