acceptodds
Under review as a conference paper at ICLR 2027

Can Omni-modal LLMs Understand Spoken Intentions?

Abstract

Omni-modal large language models (Omni-modal LLMs) can answer questions about audiovisual content, yet useful assistance also requires understanding what help the user intends to obtain. This paper introduces \ourbench for evaluating spoken intention understanding through contextual responses. The central requirement is to determine the valid request, retain its constraints, and respond according to the available evidence. We accordingly organize the benchmark around request grounding, request and context selection, assistance need assessment, request integration, and responding under uncertainty. Beyond direct requests, it examines indirect help seeking, revised requirements, ambiguous goals, and false premises. Paired scenarios vary request and evidence conditions to test whether models adjust their responses appropriately. A scenario-driven pipeline constructs scripts, generates audiovisual interactions, and defines response rubrics for language, intent, and content. Recorded videos with synthesized questions provide complementary spoken QA evaluation and training data. We further develop \ourmodel through reinforcement learning to examine transfer from explicit spoken QA to complex request conditions. Evaluation reveals substantial variation across models and tasks. Post-training improves instruct-mode performance, while uncertainty handling and selective responding remain weak. These findings establish \ourbench as a diagnostic resource for assessing spoken intention understanding and identifying priorities for model improvement.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.