GaitReader: From Large-Scale Healthy Gait Pretraining to Cross-Center Disease Recognition
Abstract
Six-degree-of-freedom (6-DoF) knee kinematics provide a non-invasive means of assessing lower-limb function and recognizing musculoskeletal disorders. However, existing learning-based approaches largely depend on disease-specific annotations, which are difficult to obtain at scale across diverse patient populations and clinical centers. Although learning from healthy movement has shown promise for gait analysis, whether healthy-only pretraining can support label-efficient knee disease recognition across centers remains insufficiently explored. To investigate this question, we introduce KineGait, a multicenter dataset containing 4,168 healthy participants for self-supervised training and validation, alongside Healthy, anterior cruciate ligament deficiency (ACLD), and knee osteoarthritis (KOA) cohorts for downstream evaluation. Building on this resource, we propose GaitReader, which organizes recordings into cycle-level gait units and learns a gait vocabulary comprising six DoF-specific codebooks. The frozen tokenizer provides discrete targets for contextual pretraining, combining masked code prediction with attribute prediction and waveform reconstruction without disease annotations. Following supervised fine-tuning, GaitReader achieves mean accuracies of 93.04% and 89.87% on fixed internal and external test sets, respectively, across five training–validation folds. Pretraining improves accuracy over the same architecture trained from scratch at every evaluated label fraction; using half of the training labels, it exceeds scratch training with all labels on both test sets. These findings support healthy gait as a source of transferable representations for label-efficient disease recognition across the evaluated centers.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.