acceptodds
Under review as a conference paper at ICLR 2027

Image-Defect Evidence for Task-Fidelity Configuration of Frozen Driving Video Generators

Abstract

Driving-video generators provide synthetic observations for offline evaluation of autonomous-driving systems, but visually plausible outputs can alter perception, tracking, and planning metrics relative to condition-aligned real observations. Inference-time controls offer a way to reduce these discrepancies without retraining, yet evaluating every configuration with the full driving stack is costly. We present a reference-based configuration workflow that uses image-defect measurements to prioritize downstream evaluation while keeping both the generator and evaluators frozen. Single-factor screening identifies controls that alter task-associated defects, and a lightweight surrogate ranks candidates within the retained search space. Each search round sends only the best-ranked candidate for full downstream evaluation, while a separate multi-task check assesses nominated configurations on held-out scenes. Experiments with the WorldDreamer renderer and a Cosmos-based pipeline on 26 InterHub scenes show that reconstruction and temporal-consistency defects are associated with perception and planning gaps in both pipelines after accounting for scene and generator effects. Held-out evaluation identifies one configuration per pipeline that reduces the composite gap while satisfying task-specific tolerance checks under the frame-level decision rule. The relative reductions are 31.24% and 21.02% for WorldDreamer and Cosmos, respectively, corresponding to absolute reductions of 0.1415 and 0.0952 on the fixed percentile-normalized scale. The WorldDreamer configuration primarily narrows the tracking gap, while the Cosmos configuration narrows both perception and tracking gaps. In a matched single-run comparison, guided search reaches a lower task gap using 41% of the GPU time recorded for random search. These results support image-defect-guided configuration of frozen generators for offline driving evaluation, with image measurements prioritizing candidates and observed task-metric gaps determining which configurations to retain.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.