Commutator-Guided Learning of Order-Aware Lie Actions in Mixed-Latent VAEs
Abstract
A key challenge for Lie-structured variational autoencoders is to learn how actions compose and reproduce their order-dependent physical effects in decoded observations. We propose a two-phase framework separating representation learning from physical-action grounding and algebraic refinement. Phase 1 (P1) combines a categorical branch for category-related content with a continuous branch for Lie transformations. We compare unsupervised training with paired-weak supervision, which uses observation pairs sharing category or transformation state. Phase 2 (P2) first grounds latent directions using action-labeled transitions while keeping the representation model frozen (P2a), then refines these directions and Lie generators toward prescribed commutator relations while retaining finite-action targets (P2b). To connect the algebraic refinement to observable accuracy, a rollout variant (P2b+) additionally matches decoded action trajectories through a frozen decoder. We evaluate latent actions, decoded observations and Lie structure on synthetic image benchmarks and controlled global rotations of real human-pose data, covering non-commuting actions and commuting controls. Paired-weak P1 improves mean categorical and geometric specialization, with trade-offs in reconstruction and FactorVAE (FVM) scores. In P2, algebraic refinement strengthens prescribed non-zero relations beyond grounding alone, while P2b+ achieves lower mean decoded endpoint error than Core P2b. Compared with the Commutative Lie Group VAE (CLG-VAE)-style strict control, which naively encourages commutativity by penalizing all basis-pair commutators, grounded refinement improves action-direction agreement across several image benchmarks, better recovers prescribed non-zero brackets, and reduces unwanted decoded order effects on the commuting benchmarks. These action-level gains coexist with competitive reconstruction and FVM scores relative to the strict control and external baselines. Results on the skeleton datasets likewise show improved recovery of non-commuting order effects and prescribed bracket directions and magnitudes. Together, these findings support grounding and refining the appropriate action relations rather than imposing blanket commutativity, while distinguishing algebraic agreement from accurate observable effects.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.