acceptodds
Under review as a conference paper at ICLR 2027

DepthFoley: Depth-Guided Video-to-Audio Generation for Spatial Consistency

Abstract

Current Video-to-Audio (V2A) approaches have achieved impressive semantic quality but often treat the auditory environment as a spatially-static and distance-agnostic plane. Neglecting fundamental distance-dependent acoustic decay, such as sound attenuation and spectral filtering, results in audio that fails to reflect spatial dynamics and shatters the immersive illusion. To address this, we propose DepthFoley, a framework that shifts the paradigm from purely data-driven fitting to a depth-guided alignment strategy. By leveraging metric depth estimation, we extract metric depth trajectories to guide a pre-trained backbone model via a Depth ControlNet. We further introduce a training-free object-wise inference strategy for multi-object scenes. To ensure spatial consistency without sacrificing generative priors, we introduce Sound Pressure Level (SPL) and Spectrally Perceptual Energy (SPE) losses. Augmented by depth-informed regularization, these objectives guide the model to internalize realistic decay trends while preserving disentangled acoustic dynamics, balancing veridical spatial realism with artistic flexibility. To evaluate this, we introduce DITC and DIGA, two novel metrics experimentally validated to align with human spatial perception. Extensive experiments demonstrate that our architecture achieves superior spatial-acoustic alignment and strong zero-shot generalizability to complex, out-of-distribution environments, establishing a new standard for spatial perceptual realism.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.