DivLAM: Latent Action Discovery in World Models through Demonstrator Diversity
Abstract
Stochastic execution makes action discovery from video ambiguous: a demonstrator's chosen command need not match the movement executed. DivLAM separates these hidden choices in two stages. Stage A learns shared next-state predictors and demonstrator-specific executed-action probabilities in a frozen visual representation. Stage B factors these probabilities into command policies and an execution channel specifying movement probabilities given each command. Building on existing minimum-volume identification theory, we prove that policies covering every ordered pair with probabilities and satisfy the required sufficient-scattering condition for at least four commands. A matched family restricted to cyclic neighbors instead yields an absorbed channel that incorporates policy mixing. Exact-input tests recover the predicted targets, while sampled counts distinguish sampling error from persistent bias. Across five fresh TwoRoom replication datasets, all-run mean channel total-variation errors after relabeling are for ordered pairs and for cyclic, versus for ignoring execution noise. This visual gap reflects both different population targets and unequal success in learning the executed basis. Discovery uses no action labels, but the frozen encoders were pretrained with actions; recovery is sensitive to encoder choice and fails on the tested PushT task. The results connect demonstrator policy coverage to channel recovery and distinguish factorization bias from errors in learning the executed basis.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.