acceptodds
Under review as a conference paper at ICLR 2027

Share Poses, Not Pixels: An Interventional Study of Shared World Models for Multi-Agent Navigation

Abstract

A navigation world model predicts what an agent will see after acting. It is natural to ask whether several agents could share one, each broadcasting its intent so that every agent's imagined future contains the others, and whether such a shared model could serve as a negotiation device, scoring one's own plans by the reactions they provoke. We test this with a conditional diffusion world model in the style of Navigation World Models, extended with a peer channel, inside a simulator in which interaction strength is a dial and a control world with exactly zero interaction exists for every scene. (i) The model learns to use a broadcast pose: a matched difference-in-differences against the control puts the broadcast at -0.014 LPIPS (95% CI [-0.017,-0.012]), five to eight times its value where nothing interacts. (ii) In closed-loop decentralised planning against neighbours of unknown disposition, reading the neighbour out of the imagined frame is no better than ignoring it (46.7% vs. 48.9% collisions) while a six-bit broadcast of its declared pose matches an oracle (2.2%), and the gap widens from 17 to 67 points as the crowd grows from 2 to 5. A ladder of oracles attributes this to field of view, the pixel readout, and the imagination, which draws a neighbour within one metre at 0.42 of its size; a readout fitted to the model's own frames does not recover it. (iii) A state head trained jointly on the same trunk does: it places the neighbour 3.2 s ahead to 0.15 m in one forward pass and learns how the neighbour reacts to the ego's action (slope 0.39 against 0.00 in the control world), a dependence the pixel objective never acquires. Yet in closed loop this head collides nearly three times as often as constant-velocity extrapolation from the same input, because what it learned is the training population's policy and the neighbour in front of it may not share it. Pixels lose the number; a learned model of others must bet on their policy; a declared intent does neither. In this setting, sharing should happen at the level of poses and intents rather than pixels.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.