CERPA: Heading-Equivariant Attention for Yaw-Invariant Few-Shot Room-Acoustic Synthesis
Abstract
State-of-the-art cross-room few-shot acoustic synthesis methods generate room impulse responses (RIRs) from sparse acoustic context and geometry conditioningf. However, their geometry conditioning often fails to respect the heading invariance of monaural acoustics: jointly applying a yaw rotation to the panoramic depth image and source–receiver poses can change the predicted RIRs. Such a change of heading reference changes only the coordinate frame, leaving the underlying acoustic characteristics unchanged; the conditional probabilistic RIR distribution or deterministic RIR output should therefore remain invariant. To address such physical inconsistency, we introduce CERPA (Cyclic-Equivariant Relative Positional Attention), a unified attention framework that explicitly builds this rotational invariance into the visual representation. CERPA modifies the attention mechanism of the geometry-conditioning ViT to produce cyclic rotation equivariant tokens, followed by an invariant readout for heading-invariant visual conditioning. When integrated into both deterministic and state-of-the-art probabilistic acoustic models, CERPA ensures yaw invariant acoustic synthesis. In both simulated and real-world few-shot RIR synthesis, our approach substantially improves heading consistency and outperforms non-equivariant baselines on key acoustic metrics.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.