RiskNav: Safety-Readable Latent World Models for Crowd Navigation
Abstract
Safe crowd navigation requires anticipating how candidate robot actions affect future human-robot interactions under partial observation. Reactive methods do not explicitly evaluate these futures, whereas predictive planners that optimize multi-agent trajectories online can become costly as crowds grow. Latent world models offer compact prediction, but their representations alone do not provide explicit safety quantities for risk-aware action selection. We present RiskNav, a risk-supervised, action-conditioned latent world model for crowd navigation. For each candidate action sequence, RiskNav rolls out future latent representations conditioned on the observation history and local waypoint. This latent rollout avoids explicitly predicting individual pedestrian trajectories or optimizing multi-agent trajectories online. A risk evaluator predicts static and dynamic safety margins and collision probabilities from each representation, with gradients from the resulting risk loss propagating through the predictor into the encoder. The planner incorporates these predicted safety quantities into a horizon-wise risk term for receding-horizon planning. Across indoor, open, and out-of-distribution settings, RiskNav achieves the highest strict success rates and the fewest robot-responsible contacts among nine baselines, with strict success rates of 0.76, 0.57, and 0.50, respectively. Replanning stays near 42ms as the crowd grows from 1 to 21 pedestrians. Ablations show that RiskNav's advantage comes from coupling safety-supervised latent prediction with risk-aware action selection, as strict success falls from 0.750 to 0.350 without representation shaping and to 0.125 without the planning-time risk term. RiskNav makes action-conditioned future risk directly usable for planning, improving both safety and success in crowd navigation.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.