A Gaze-Free Human-Inspired Object Detection Explainer via Evidence-Based Fixations
Abstract
Explanations for object detectors are usually evaluated against detector behavior, leaving their alignment with human attention unclear. Human-inspired explainers largely target image classification, while human-supervised methods learn from gaze and cannot use it as an independent reference. We introduce HiOdex (Human-inspired Object Detection Explainer), the first post-hoc, human-inspired explainer that produces object-specific attributions for object detectors without gaze supervision. HiOdex derives model-generated fixations from activation- gradient evidence, applies inhibition of return, and uses their envelope to guide class activation mapping (CAM). Local evidence and multi-layer fusion preserve spatial detail in a single forward and backward pass. We compare HiOdex with eight explainers across two detectors and two benchmarks using model-oriented metrics, then independently assess human-gaze alignment on two gaze datasets using image-level saliency metrics and object-level measures. HiOdex achieves the highest Over-All score among all evaluated CAM-based and human-inspired baselines in all four detector and benchmark settings. It comes within 0.0217 of the strongest perturbation-based result at roughly 1% of the computational cost, achieves the best or tied-best AUC-Judd in every gaze setting, and ranks in the top two for 8 of 12 evaluations using image-level gaze metrics. These results show detector-derived fixations can yield efficient explanations that perform strongly under both model- and human-centered evaluation without gaze supervision.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.