acceptodds
Under review as a conference paper at ICLR 2027

HumanSkill-4D: Training-Free Language Querying of Animatable Gaussian Avatars in Space and Time

Abstract

We introduce avatar referring segmentation, which uses a language query to select a region of an animatable Gaussian avatar and, when the query specifies a pose or motion, the matching intervals in its driving pose sequence. Such queries can support local editing and motion inspection of digital humans. Adapting existing language Gaussian splatting methods to avatars poses two challenges: adjacent parts and garments are hard to separate and may be referred to indirectly (region ambiguity), while fine-grained motions can look alike in RGB (motion ambiguity). To address these challenges, we propose HumanSkill-4D, a training-free framework with a multimodal large language model (MLLM). For region ambiguity, it builds a semantic memory linking the avatar’s parts and garments to their persistent Gaussian regions and appearance descriptions, from which the MLLM selects target categories, even when referred to by appearance, relation, or function. For motion ambiguity, it stores joint trajectories in a motion memory, and the MLLM composes geometric operators into a motion program that a deterministic executor runs on this memory to find matching intervals. On three human-annotated benchmarks, HumanSkill-4D outperforms prior methods in 3D and 4D referring segmentation and also in part segmentation.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.