acceptodds
Under review as a conference paper at ICLR 2027

Where Will It Fail? Detector-Native Concept Blindspots for Forecasting Errors

Abstract

Object detection underpins numerous safety-critical applications such as autonomous driving and embodied AI, yet even the best detectors still make perception errors that pose serious risks. Understanding detector failure modes is therefore important for diagnosing model vulnerabilities, assessing risk, and guiding improvements. A detector’s intrinsic concepts may play a fundamental and general role in determining its errors, but explicitly extracting and quantifying such concept-level vulnerabilities remains challenging. We introduce a framework to discover, validate, and leverage detector-native concept blindspots for failure diagnosis and risk screening. We develop DetSAE, our task-relevant sparse autoencoder, to extract sparse concepts from a fixed detector's object-level representations while preserving its classification and localization behavior. We then generate 20 style-transferred versions of the original dataset and track objects to identify blindspots associated with statistically supported excess detection degradation. Experiments demonstrate that suppressing excessive blindspot activations improves frozen-head classification across detector backbones, providing causal evidence of their contribution to errors. Using the same frozen blindspot representation, our separate false-negative and false-positive risk predictors achieve the best performance among compared methods across diverse benchmarks. These results show that seemingly isolated detection failures can share concept-level weaknesses that support both model interpretation and actionable risk assessment. Our code and generated datasets are publicly available on GitHub.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.