Consistency Is Not Identification: What Reuse Reveals About Latent Actions
Abstract
Latent-action models learn action representations from unlabeled video by encoding observed transitions into codes. Many methods shape these representations using algebraic constraints or cycle-consistency checks that re-encode a model's own predictions. However, we prove that a model can fit the observed trajectories and satisfy these constraints yet predict the wrong effect when an action is reused in a new state. In this sense, consistency is not identification. We then show which data resolve this ambiguity. For reversible actions on a connected finite state graph, repeated-command trajectories and trials that re-execute a command after an intervening one together uniquely determine the set of true action operators. To evaluate learned actions in this setting, we further show that, for deterministic models with as many codes as commands, the distance between learned and true operators is bounded by local reconstruction error plus reuse error at independently sampled targets. In controlled experiments, training on re-execution trials learns the true state-transition operators on every restart, whereas training only on repeated-command data can yield consistent but incorrect operators. In LAPO models across Procgen environments, codes inferred at the target select the correct outcome in of eligible pairs, compared with for codes transferred from another state on the same pairs. These results show that the gap between reconstructing observed transitions and reusing actions in new states also appears in latent-action models trained on video.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.