Evaluating Any-to-Any Instruction Following with Omni-modal Instruction
Abstract
Any-to-any models are expected to follow user instructions and respond in multiple modalities. In this setting, instruction following is not limited to satisfying content or format constraints, but also requires understanding how the instruction is conveyed. To extensively assess this capability of any-to-any systems, we introduce a novel benchmark, Omni-Instruct, which evaluates instruction following in any-to-any models with various context, instruction and output modalities across text, image, and audio. Using our proposed framework, we extensively evaluate recent any-to-any models, decoupling the output-modality routing from instruction compliance. From our study, we discover that 1) routing failure is the dominant factor for most apparent performance drops, 2) even when correctly routed, models show distinct compliance failure patterns, and 3) models remain vulnerable to omnimodal instructions, struggling to understand and follow requests conveyed through images or speech.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.