Learning Target-Centered Latent Actions through Predictive Comparisons
Abstract
When several agents act together, observations reveal only their shared transition, making it difficult to learn a latent action for any one agent. A representation trained to explain the full transition may mix the target agent with its partners, while simply separating agents can discard how the target’s effect depends on them. We introduce a target-centered approach based on controlled predictive comparisons. A learned dynamics model makes two predictions from the same history and the same partner representations, changing only the target representation. Their difference isolates what changes with the target while holding the surrounding context fixed. Repeating this comparison across partner settings reveals how the target’s predicted effect varies with individual partners and their combinations. We compress these structured differences into a latent action, train a history-only predictor to recover it before acting, and use a small action-labelled dataset only to ground it in the target action space. Our analysis characterizes what these comparisons remove, what partner-dependent information they retain, and when the resulting representation can be recovered from history. Experiments on particle coordination, multi-agent locomotion, and visual observations show that the learned representations retain target-action and partner-dependent information and improve latent-action control under limited action supervision.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.