Temporal CAMML: Confidence-Aware Temporal Motion Masking for Self-Supervised Autism Behavior Representation Learning
Abstract
Video-based analysis of autism behavior can help with the objective assessment of repetitive and stereotypical behavior. However, existing skeleton-based approaches are mostly dependent on fully supervised learning and large annotated datasets, which are hard to obtain for autism research. To address this, we propose Temporal CAMML, a self-supervised representation learning framework for autism behavior analysis based on confidence-aware temporal motion masking. In the proposed framework, the 17-keypoint pose sequences from videos are first extracted by YOLO11-Pose. Then each behavior event is represented by a normalized temporal pose tensor. Temporal CAMML differs from traditional masking techniques that remove individual joints by identifying low-confidence skeletal observations and masking connected temporal segments of motion, thus encouraging the model to learn robust motion dynamics. A transformer encoder is then trained to reconstruct the masked trajectories, allowing for self-supervised representation learning without behavior labels. We then evaluate the learned representations with a linear probe on downstream behavior classification . Experiments on the self-stimulatory behavior data with 102 annotated events of arm flapping, head banging and spinning demonstrate the effectiveness of the proposed approach. Temporal CAMML achieves 61.9% classification accuracy, outperforming a supervised Transformer baseline (52.38%) and a confidence-aware masking variant without temporal motion modeling (52.38%). Our results demonstrate that confidence-aware temporal masking facilitates learning more discriminative pose representations for autism behavior analysis in low-resource settings.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.