acceptodds
Under review as a conference paper at ICLR 2027

EdgeGaze: Efficient End-to-End Multi-Person Gaze Target Detection

Abstract

Gaze target detection on edge devices requires identifying where every person in a scene is looking under tight memory and latency budgets. Existing two-stage methods rely on external head detection and repeated per-person gaze estimation, while recent end-to-end models can grow to hundreds of millions of parameters. We introduce EdgeGaze, a compact end-to-end model built on EdgeCrafter that jointly detects multiple people and predicts their gaze targets in a single forward pass, with computation nearly independent of scene crowding. Gaze is strongly shaped by scene structure: people often look at other people, salient regions, and meaningful objects. We exploit this through Scene Prior Guidance, using saliency, object, and head priors to learn a Prior-Guided Scene Representation (PGSR) during training. PGSR is fused with person queries for gaze prediction, while the external prior models are discarded at inference. We further introduce a Unified Gaze Distribution (UGD), which places in-frame gaze locations and the out-of-frame outcome in a single normalized prediction space. EdgeGaze achieves state-of-the-art performance on GazeFollow and VideoAttentionTarget, with strong zero-shot transfer to GOO-Real and ChildPlay. On GazeFollow, EdgeGaze-M achieves 0.960 AUC and 0.798 mAP with only 26.7M parameters and 4.9 ms latency, compared with 0.929 AUC and 0.626 mAP for the 862M-parameter GazeHTA. EdgeGaze-N further reduces the model to 8.5M parameters and 3.1 ms. These results show that accurate multi-person gaze target detection can be compact, fast, and fully end-to-end, making it well suited for edge deployment. Code and models will be released; demo: https://anonymous-edgegaze.pages.dev/

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.