MuscleSight: A Benchmark for Ultrasound Representations of Muscle State in Manual Tasks
Abstract
Arm-mounted ultrasound images the muscles that move the hand, including deep muscles that surface electromyography cannot isolate. Existing models report strong gesture and pose decoding, but are evaluated mostly within participants and sessions. We introduce MuscleSight, a dataset and benchmark for testing whether ultrasound-based hand decoding generalizes: 37.7 hours from 33 participants across 113 sessions and 37 manual tasks, including isolated motions, static postures, and interactions with 72 distinct objects. Synchronized multi-view videos provide audited hand-pose and gesture labels. Our benchmark evaluates frame- and clip-based gesture classification, pose regression, and pose tracking within sessions and on held-out sessions and participants. Within sessions, a 0.28 M-parameter CNN reaches 89.5% gesture accuracy (majority baseline: 21.0%). On held-out sessions and participants, however, gesture classifiers fail to beat the majority baseline, pose regressors improve over a constant mean pose by at most 1.8°, and only initialized pose tracking transfers reliably. Learned representations strongly encode recording session, revealing substantial session-specific structure. Few-shot calibration on each new session improves every calibrated model, while simply adding training participants does not: a calibrated pose model trained on 2 participants outperforms an uncalibrated model trained on 28. Data, labels, splits, and code will be released.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.