acceptodds
Under review as a conference paper at ICLR 2027

LUNAR: Language-Grounded Latent Reasoning for Text-to-Motion Generation

Abstract

We propose LUNAR, a structured textual–latent reasoning framework built on a core principle: use language to scope latent reasoning, and motion to ground it. LUNAR organizes complex motion into atomic motions and represents each as a paired textual–latent reasoning unit. The textual block specifies the action intent and how it is realized across the body, defining the semantic scope of the latent block. The latent block is grounded in the corresponding motion segment through a pretrained text–motion space. After completing the reasoning trajectory, LUNAR predicts motion tokens that can be decoded into 3D motion by a frozen motion codec. We train LUNAR with progressive supervised fine-tuning followed by trajectory-level group relative policy optimization. Experiments show that LUNAR establishes state-of-the-art in-domain (ID) text–motion alignment on HumanML3D, achieving a Top-3 R-Precision of 0.838 and an MM-Dist of 2.714. In out-of-domain (OOD) evaluation across compositional, paraphrastic, and cross-dataset shifts, LUNAR consistently outperforms all compared baselines, improving overall R@3 from 0.517 to 0.593 in the T2M space and from 0.606 to 0.696 in the independent LaMP space. Our code will be publicly available upon acceptance.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.