acceptodds
Under review as a conference paper at ICLR 2027

OmniGym: A Closed-Loop Evaluation Framework for Full-Duplex Omni Models

Abstract

Real-time omni models support continuous audio-visual interaction, where users can interrupt, revise requests, and resume earlier tasks. Evaluating these capabilities requires more than replaying predefined input trajectories: full-duplex models should be exercised inside evolving interactions, where simulated user behavior adapts to model responses while perceptual and runtime conditions change over time. We introduce OmniGym, a programmable framework for closed-loop evaluation of full-duplex omni models. OmniGym represents each test case as an executable interaction episode, using an event-driven user simulator and a controllable environment emulator to realize adaptive user behavior and environment shifts along a shared timeline. A timeline DSL captures the event dependencies and evaluation objectives of each episode, allowing the same interaction logic to be reproducibly executed across models and conditions. Using OmniGym, we construct 200 scenarios spanning seven categories of interaction and perceptual distractors and their combinations, and evaluate four systems from the Qwen, Gemini, and GLM families. Our evaluation reveals substantial variation across distractor types and severity groups: different models excel under different interaction and perceptual conditions, while exhibiting markedly different trade-offs among semantic correctness, interaction behavior, and response and interruption latency. These results show that full-duplex capability cannot be characterized by a single metric or fixed interaction trajectory, motivating closed-loop evaluation that exercises how models behave as users and environments evolve.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.