AutoClAM: End-to-End Differentiable Deep Clustering for Unsupervised Action Localization
Abstract
Unsupervised Temporal Action Localization (UTAL) aims to discover and segment action instances within untrimmed videos without ground-truth annotations. Existing clustering-based pipelines typically rely on disjoint two-stage frameworks that apply discrete, non-differentiable clustering (e.g., k-means) to frozen representations. Such discrete assignments inherently sever the computational graph, preventing gradient propagation from the localization objective back to the feature encoder. To overcome this limitation, we propose a novel end-to-end differentiable deep clustering framework for skeleton-based UTAL that seamlessly unifies representation learning and temporal boundary discovery. We formulate temporal clustering as a continuous energy-minimization process by coupling a Spatio-Temporal Graph Encoder with Continuous Local Associative Memories (ClAM). During training, our model jointly optimizes the feature manifold and continuous action prototypes via a composite differentiable objective combining masked pattern reconstruction with continuous soft-relaxations for intra-cluster compactness and inter-cluster separation. Furthermore, to eliminate the field's rigid reliance on pre-defined cluster counts, we introduce the Automated Continuous Local Associative Memory (AutoClAM) inference module. Operating strictly at inference time, AutoClAM employs recursive splitting guided by the Bayesian Information Criterion (BIC) to dynamically discover the optimal number of distinct underlying motions within each unseen sequence. Extensive evaluations across three complex skeleton benchmarks (HuGaDB, BABEL, and LARa) demonstrate that our framework consistently outperforms state-of-the-art methods, including recent Optimal Transport baselines, especially at strict overlap thresholds (F1@50). We further show that our continuous formulation exhibits linear computational complexity O(T * k * D) and empirically generalizes to dense RGB video features. Code will be released upon acceptance.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.