PragyaMD: Parameter-Efficient Adaptation of Humanoid Motion Generators to Indic Languages
Abstract
English-centric text encoders can obscure distinctions between Indic instructions before those instructions reach a humanoid motion generator. We address this bot- tleneck with PragyaMD, a 4.7M-parameter adapter that extends a pretrained, frozen OMG motion diffusion model to Hindi, Bengali, Tamil, and Telugu, with- out inference-time translation or generator retraining. PragyaMD maps representations from a frozen multilingual encoder into the gen- erator’s conditioning space. A single shared adapter is trained on Unitree G1- format trajectories using a motion-pivot objective comprising (i) motion-denoising supervision, which couples each instruction to its corresponding observed motion continuation; (ii) cross-language agreement, which encourages compatible con- ditioning across language views of the same motion; and (iii) English-pathway regularization, which weakly anchors the learned interface to the model’s origi- nal English conditioning pathway. Motion thus provides a shared training target across languages; every component except the adapter remains frozen. On 1,824 leakage-controlled motion groups, PragyaMD improves English- mediated top-1 retrieval from 13.3–14.8% with direct Indic input to 27.5–28.3% across the four evaluated Indic languages, compared with 26.2–27.2% for an In- dicTrans2 translate-then-generate baseline. The translation baseline has 2.4–2.6× PragyaMD’s Fr´ echet distance in the evaluator’s motion space. A matched English- adapter control shows that the Fr´ echet improvement arises from replacing the conditioning pathway, not from Indic-language adaptation itself. Text-encoder- free interventions reveal consistent forward/backward and speed responses but weak left/right separation; cross-language tests show consistent responses across languages but expose sensitivity to instruction wording. These findings sup- port parameter-efficient language-interface adaptation for extending pretrained humanoid motion models, while showing that retrieval improvements alone do not establish reliable attribute-level instruction grounding.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.