Preserve, Then Bound: Conditionally Informative Acoustic Context for Few-Shot ATC Call-Sign Spotting
Abstract
Acoustic context in few-shot keyword spotting is neither uniformly irrelevant nor consistently reliable. Speaker and channel cues can disambiguate similar lexical targets while a local source association persists, yet become misleading when that association changes. We introduce , a framework that treats these cues as conditionally informative context while retaining lexical content as the primary evidence. summarizes class-stable lexical content with class prototypes and preserves transient acoustic context in an instance-level support memory. A pairwise relation network derives a candidate-wise contextual applicability cue for each query-candidate pair, using joint class- and source-agreement supervision during training without requiring source labels at inference. This applicability cue and margin-based lexical uncertainty jointly modulate a bounded contextual residual, allowing contextual evidence to refine uncertain lexical comparisons without granting it unrestricted decision authority. The bounded update preserves the lexical top-1 prediction whenever the lexical margin exceeds the maximum possible context-induced score difference. We evaluate on air traffic control call-sign spotting in ATCO2 and UWB-ATCC under class-disjoint 5-way 1-shot and 5-shot protocols with clean and acoustically degraded queries. Controlled ablations examine the contributions of instance memory, source-context supervision, contextual applicability, and lexical-uncertainty conditioning. By separating contextual applicability from decision authority, provides a strong lexically anchored recognizer and extends it with transient acoustic evidence, refining ambiguous comparisons while preserving sufficiently decisive lexical predictions.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.