acceptodds
Under review as a conference paper at ICLR 2027

Query Anytime: Find When Before Where for Online 4D Point Cloud Segmentation

Abstract

Query-driven 4D point cloud segmentation aims to segment target instances in dynamic 3D sequences based on natural language queries. Existing methods largely follow a query-first offline setting, where the query is given in advance and the input sequence is temporally bounded around the queried target. However, in real-world interactions, queries may arrive at arbitrary times while 3D observations continue to evolve. In such an online scenario, the query and its relevant observations are no longer temporally aligned in advance, requiring the model to establish their correspondence dynamically. We summarize this challenge as “when before where”: the model must first identify when the queried target is present in the evolving sequence, before determining where it should be segmented. Based on this principle, we propose QueryAnytime. For when, a Searchable Evidence Tokenization (SET) module is proposed to build a compact, query-independent evidence space for efficient anytime query grounding without revisiting dense history. For where, a TrackState module is designed to update a structured semantic-geometric target state, which maintains target identity and spatial continuity for instance-consistent segmentation over time. We further introduce Online4DKITTI, a dedicated benchmark for online query-driven 4D segmentation with arbitrary query timing and fine-grained 3D instance annotations. Experiments demonstrate that QueryAnytime substantially outperforms existing baselines in segmentation accuracy while maintaining acceptable inference latency. Code and benchmark will be released at: https://anonymous.4open.science/r/QueryAnytime-E979/

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.