acceptodds
Under review as a conference paper at ICLR 2027

AnchorSLU: Cross-Lingual Semantic Anchoring for Low-Resource Speech Understanding

Abstract

Multilingual speech foundation models have advanced rapidly, yet large resource-dependent performance gaps persist when target-language supervision is limited. Standard cross-entropy (CE) provides only *class-level supervision*: semantically different utterances within the same class receive the same target, leaving exact sentence-level cross-lingual correspondence unused. We introduce ANCHORSLU, which converts this correspondence into *representation-level supervision*. For each target utterance, we synthesize sentence-matched reference speech in eight high-resource languages, encode these reference views offline with a frozen speech backbone, and aggregate selected-layer representations into sentence-specific semantic anchors. During adaptation, the model learns from natural target speech using target-only CE and multi-layer anchor alignment; inference requires only the adapted target model. Across 20 low-resource SIB-Fleurs languages, ANCHORSLU improves Qwen2.5-Omni-7B over matched target-only CE by **4.39** points in Accuracy and **5.04** points in Macro-F1. On 60-intent Speech-MASSIVE, it improves Accuracy by **3.19** points and Macro-F1 by **3.52** points under substantially richer target supervision. The benefit further extends to a substantially larger speech backbone. Controlled comparisons show that removing exact sentence correspondence largely eliminates the benefit even when class identity is preserved, while direct TTS augmentation and output-level distillation remain weaker. Together, these results establish exact sentence-level cross-lingual correspondence as an effective source of representation-level supervision beyond class labels, suggesting that sentence-matched multilingual supervision may also benefit broader speech understanding settings where such correspondence is available.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.