When to Trust Context: A Causal Evidence Gate for Occlusion-Aware Object Detection
Abstract
Transformer-based object detectors aggregate scene context through cross-attention causing query representations to entangle an object's direct visual evidence with contextual co-occurrence priors. Traditionally treated as distinct challenges, we demonstrate that this contextual entanglement underpins two opposing failure modes: (i) co-occurrence bias where models over-rely on spatial surroundings, impairing generalization under distribution shifts and long-tail categories; and (ii) partial observability where context serves as the sole signal for recovering heavily occluded objects. We unify these phenomena through a single causal framework and introduce Evidence-Gated Backdoor Adjustment (EG-BA) for Open-Set Object Detection (OSOD). EG-BA employs a learned per-query evidence score to dynamically interpolate between two regimes: when visual evidence is reliable, it executes a backdoor adjustment to strip confounding context; when evidence is degraded, it leverages context for identity recovery. To ensure the internal mechanism behaves in the exact way, we regularize the latent space via a variational information bottleneck and propose the Gate-Occlusion Calibration Error (GOCE) a diagnostic evaluation tool to understand the gate aligns with true physical occlusion. We investigated that EG-BA improves calibration and unknown-class recall over a vanilla Deformable DETR baseline model on Pascal VOC to MS-COCO dataset and its evidence score tracks synthetic occlusion severity when properly regularized, without this regularization the same information-bottleneck term drives the gate to a collapsed input-independent constant state instead, considered as a failure mode, we report as a central finding of independent interest. We further find that this occlusion-robustness gain does not transfer to real, naturally occurring occlusion and that the mechanism remains steady for co-occurrence-driven performance degradation, the bias it is designed to remove.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.