acceptodds
Under review as a conference paper at ICLR 2027

Evaluating Any-to-Any Instruction Following with Omni-modal Instruction

Abstract

Any-to-any models are expected to follow user instructions and respond in multiple modalities. In this setting, instruction following is not limited to satisfying content or format constraints, but also requires understanding how the instruction is conveyed. To extensively assess this capability of any-to-any systems, we introduce a novel benchmark, Omni-Instruct, which evaluates instruction following in any-to-any models with various context, instruction and output modalities across text, image, and audio. Using our proposed framework, we extensively evaluate recent any-to-any models, decoupling the output-modality routing from instruction compliance. From our study, we discover that 1) routing failure is the dominant factor for most apparent performance drops, 2) even when correctly routed, models show distinct compliance failure patterns, and 3) models remain vulnerable to omnimodal instructions, struggling to understand and follow requests conveyed through images or speech.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.