acceptodds
Under review as a conference paper at ICLR 2027

Second-Order Motion Representations: Grassmannian Differential Kinematics for Unsupervised Hand-Object Interaction Segmentation

Abstract

Representations of human motion are typically built from appearance, spatial configuration, or first-order displacement. We show that what they omit is second-order structure: the rate at which motion direction and speed change. This is not recoverable from zero- or first-order descriptors, yet it is what distinguishes execution modes sharing objects, locations, and appearance - tightening versus loosening, moving versus pouring. We propose a representation that makes this structure explicit. Short windows of 3D scene flow are lifted to points on the Grassmann manifold, yielding a subspace trajectory whose covariant derivatives expose manifold acceleration and direction change as independent, coordinate-free descriptors of how motion evolves. Recovering them from noisy observations poses a representational trade-off: differentiation amplifies noise, while temporal averaging suppresses the transitions that constitute the signal. We resolve this by combining localized Fr\'echet means with a preemptive jump-diffusion process that halts integration across physical discontinuities.We analyse the resulting representation formally, deriving a discrete curvature error bound under stated assumptions and establishing the necessity of second-order features for a constructed class of kinematically ambiguous boundaries. The representation enables causal and training-free boundary detection. Detected segments are subsequently represented by fixed-dimensional spatial path-signature descriptors and grouped by histogram-intersection similarity with complete-linkage clustering, allowing recurring action instances to receive common discovered labels. We test the representation on unsupervised temporal segmentation. When all methods are given the same 3D scene flow, ours performs best, showing that the gains come from the representation and not from the use of 3D input, and it improves over both zero-order models and recent segmentation methods. We validate our error bounds empirically under severe noise and in realistic settings with estimated scene flow, evaluating on OakInk2, HOT3D, and EPIC-KITCHENS. Accuracy degrades predictably as the kinematic signal-to-noise ratio of the input flow falls, showing that our method operates close to the information limit of the input rather than being limited by the representation itself.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.