acceptodds
Under review as a conference paper at ICLR 2027

Learning Predictive Scene Representations from Raw 3D Gaussian Splat Parameters

Abstract

Gaussian splatting has become a common output of scene capture, and representation learning on splats has mainly distilled image features or reconstructed Gaussian parameters. Inferring what a scene hides, or anticipating what an edit will change, requires a representation that is predictable from context. We learn such representations directly from raw Gaussian parameters by predicting the latent representations of masked scene regions from their visible surroundings; images supply training targets only, so inference needs neither photographs nor rendering. Measuring recognition and prediction separately reveals that they are distinct and that image supervision trades one for the other: stronger supervision improves recognition while steadily lowering predictability, and features lifted from photographs, although better recognizers, are far less predictable than ours under matched predictors. Adding relative spatial position to the predictor makes its predictions depend on how a scene is arranged. These representations enable two capabilities: predicted embeddings retrieve real Gaussian groups from other scenes to fill hidden regions, and an action-conditioned latent world model, trained separately on synthesized edits, predicts how a scene's representation changes when objects are removed, moved, or added. Its removal predictions improve steadily with more training scenes, and searching over its predictions recovers the edits that produced a goal scene far more often than chance. Captured splats can thus be used to predict, not only to render.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.