Are Latent Actions Transferable Across Visual Contexts?
Abstract
Reusing behavioral knowledge from unlabeled videos in new scenes requires transferable action representations. Latent action models (LAMs) learn representations from visual transitions to predict future states, but may entangle action information with source visual context. Existing benchmarks lack a common protocol for evaluating cross-context latent-action transfer. We introduce the **L**atent **A**ction **T**ransfer Benchmark (**LATBench**), a unified protocol for evaluating action transfer across visual contexts. LATBench extracts an action representation from a source video and applies it to a different target video: the source provides the action, while the target provides the visual context. It evaluates (1) action-effect fidelity (whether the source action effect is reproduced), (2) target-context preservation (whether the target object and scene are preserved), and (3) source-content leakage (whether source appearance is copied into the generated video). We first assess how well existing LAMs transfer across visual contexts and examine how representation design affects transferability. We then investigate the balance between action information and source content in learned representations and test the limits of transfer under longer rollouts, camera motion, and viewpoint changes. Our experiments reveal a key insight: effective transfer requires suppressing source content while retaining action information that supports state changes in the target context; lower leakage alone is insufficient.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.