acceptodds
Under review as a conference paper at ICLR 2027

Orch3D: Global Compute Orchestration for Multi-Stage 3D Generation

Abstract

Modern 3D generators are increasingly multi-stage, making efficient orchestration of their growing inference computation critical for practical deployment. However, existing acceleration methods largely optimize individual stages or operators in isolation, overlooking dependencies across both the generative pipeline and sparse execution. We find that stage-wise optima are non-compositional: changing an upstream schedule can reverse downstream schedule preferences, showing that computation must be optimized over the complete generation process. Based on this finding, we introduce Orch3D, a global compute orchestration framework built on two complementary components: Pareto Compute Optimizer (PaCO), which jointly optimizes cross-stage compute allocation and within-stage sampling schedules through allocation-conditioned Pareto search, and Packed Sparse Window Attention (PSWA), a structure-preserving sparse execution scheme that converts irregular window partitions into reusable packed attention schedules while preserving the original attention semantics. PaCO determines where and how much computation should be spent, while PSWA reuses window structure and cached boundaries to reduce the cost of retained computation. Under an equal 90-minute search budget, complete-policy joint search attains feasible deployment policies in every evaluated search seed, whereas the evaluated factorized procedure does not; frozen joint policies also improve reference-model fidelity on held-out scenes. PSWA adds semantics-preserving execution acceleration; together Orch3D reaches 1.895×–3.776× end-to-end speedup after transferring frozen policies to SA3DAO without retuning.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.