PACE: A Reinforcement-Learning Agent for Accelerated 3D Gaussian Splatting
Abstract
Accelerating 3D Gaussian Splatting (3DGS) training is a central open problem for real-world deployment: per-scene optimization still costs minutes of GPU time, which highly depends on a fixed, hand-tuned schedule of densification, pruning, opacity resets, and learning-rate decay. Heuristic alternatives require significant domain expertise and manual tuning, while existing learned controllers are trained online per scene, increasing training time. To address these challenges, we propose Policy for Adaptive Compute-Efficient (PACE) 3DGS training, a reinforcement-learning agent that accelerates training and generalizes across scenes and 3DGS training backends. We cast 3DGS training as a block-wise Markov decision process over a pluggable training backend. At each 3DGS training block, the agent observes the optimization state (a compact summary of training progress) and selects controls (Gaussian thresholds, learning rates, etc.) for the next block. PACE uses a lightweight MLP trained offline on a small set of 3D scenes, with a time-value reward that favors achieving higher quality more quickly. At deployment, PACE only requires feedforward evaluation of the trained policy. Evaluated on four representative backends (3DGS, Faster-GS, DashGaussian, and LeGS) across different scene types, PACE's acceleration has a consistent mechanism in the Gaussians optimization: where optimizing a Gaussian is expensive (e.g., 3DGS, DashGaussian), it densifies sparingly and freezes growth early, matching the fixed schedule's quality with half the primitives; where it is cheap (Faster-GS), it densifies aggressively, converting capacity into quality during early training stages. The result is up to faster training (e.g., to seconds to reach dB on 3DGS), faster than the fixed schedule in all scenebackend cases, up to smaller models, up to lower peak training VRAM, and up to faster rendering. The learned schedules also transfer across backends without retraining in eight of the nine policy–backend pairs we evaluate, retaining –, and on LeGS reaching . The exception is predicted by the discovered mechanism, where a policy trained such that compaction is desired (DashGaussian) underperforms on the one backend where growth is optimal (Faster-GS).
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.