FRAME: A Fully Automated Pipeline for Scalable Multimodal Body Language Datasets
Abstract
Large-scale multimodal datasets are essential for training representations that capture non-verbal communication, yet most are limited by the cost of manual labeling and remain too small to support large-scale models. We introduce FRAME (Fully-automated Retrieval and Annotation of Motion and Expression), a fully automated pipeline for constructing large-scale multimodal datasets from natural dialogue videos, and use it to release POSTURE (Pose, Speech, and Talking-face Utterances for Representation Evaluation), a dataset of 156,020 movie clips with synchronized subtitles, facial expression frames, and body pose sequences. From segmentation to pose and face extraction, every stage of the pipeline runs without human input, unlike most existing multimodal datasets, which rely on manual annotation and consequently reach only a fraction of this scale. We show that this scale translates directly into better representations. Averaged across a zoo of 15 loss functions and pose encoder architectures, models pretrained on POSTURE achieve 6.8× the cross-dataset transfer R@1 of models pretrained on any other single comparable dataset (CAER, MELD, CMU-MOSEI, DFEW) at full size. This advantage does not come at the cost of quality: at matched dataset size, the advantage falls to near-parity (0.98×), confirming that scale, not source-specific quality, drives POSTURE’s advantage. Because construction requires no manual annotation, the same pipeline can be given more source footage to produce proportionally larger datasets at low marginal cost, making POSTURE inherently scalable. We release the dataset, the full extraction pipeline, and evaluation code to support future work on scalable multimodal representation learning. Code for the dataset pipeline can be found through this link: here
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.