EviScope: Evidence-Conditioned Sampling for Long-Range 3D Detection
Abstract
As autonomous driving extends to longer ranges, reliable 3D detection depends on using multiple sensing modalities well. Camera–radar fusion, pairing visual semantics with radar's reach, is an important direction. Query-based camera–radar detectors sample features around query positions. On MAN TruckScenes, recall varies substantially across object-level radar-return counts within the same range band. This motivates testing local radar evidence as an additional conditioning signal. We introduce , evidence-conditioned query sampling: each query adapts its spatial extent and temporal weights to local evidence. The current frame's returns form a non-learned bird's-eye evidence field, which each query reads to predict a sampling radius and frame weights. It adds parameters, of the detector, with no latency increase in our timing protocol. The module-free recipe exceeds the published mAP values in our comparison. At eight epochs, single runs reach / mAP / NDS without the module and / with it. Across seven matched four-epoch pairs, the module changes NDS by and mAP by (mean SE), leaving its incremental benefit unresolved. An exploratory subset gain fails an internally pre-specified test on two further seeds. The sampling radius tracks local evidence in all four probed runs, yet a large penalty from replacing temporal weights also occurs in a control lacking query-local evidence. These results distinguish learned allocation and inference-time dependence from demonstrated detection gains.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.