acceptodds
Under review as a conference paper at ICLR 2027

4DSuperG: Identity-Aware Super-Gaussians for Open-Vocabulary 4D Scene Understanding

Abstract

Open-vocabulary understanding of dynamic scenes requires text-selected objects to retain their identities as they move and become temporarily invisible. Existing 4D language fields distill per-Gaussian features through alpha-blended rendering, leaving primitive-level semantics ambiguous and often requiring feature compression. Identity-based language association addresses these issues in static scenes, but dynamic querying also requires the selected primitives to move coherently. We introduce 4DSuperG, which groups Gaussians into motion-coherent, identity-aware Super-Gaussians (SuperGs), making identity learning and motion prediction share the same unit. Each SuperG retrieves a full-dimensional CLIP feature through its object identity, without language-feature distillation. A query selects SuperGs once in canonical space, and their motion carries the selection through time and under occlusion. On HyperNeRF and Neu3D, 4DSuperG achieves state-of-the-art primitive-level open-vocabulary segmentation and competitive pixel-level performance, and supports zero-shot language-driven spatio-temporal editing.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.