acceptodds
Under review as a conference paper at ICLR 2027

Sequential Disentangled Representation Learning by Dual-Transformer Subtraction

Abstract

Sequential disentanglement, which decomposes temporal observations into time-invariant (static) and time-varying (dynamic) factors, is fundamental to interpretable representation learning. We present DTS (Disentanglement by Dual-Transformer Subtraction), a simple framework that combines an architectural inductive bias for static–dynamic separation. DTS employs a hierarchical U-Net of dual-stream Transformers: at each level, a static stream without positional encoding runs in parallel with a position-encoded full stream, and their difference is a dynamic residual feeding the next level. We show that each level's static code is permutation-invariant by construction and that, under a weak-dependence assumption, its mean-pooled estimate concentrates at rate . With this architecture, a compact objective suffices: reconstruction with an optional second pass and at most one covariance regularizer. On six video, speech, and time-series benchmarks, DTS achieves the competitive results with the strongest baselines, fully unsupervised and modality-agnostic.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.