CodeOmni: A Code-Driven Harness for Extensible Omni-Modal Reasoning
Abstract
The code-as-action/tool paradigm has been widely explored in prior work. However, existing approaches either underutilize the expressiveness of code or remain tailored to a single scenario. To address these limitations, we propose CodeOmni, a code-driven harness that uses executable programs to orchestrate interactions between models and harness components, providing a unified execution substrate for image, video, audio, and audio-visual tasks. In CodeOmni, frozen models generate Python programs that jointly support tool orchestration, control flow, exact computation, and task-local state management. A variable-mediated workspace further stores intermediate artifacts outside the main conversational context and exposes them only when needed, reducing unnecessary context accumulation. New tasks can be incorporated through lightweight runtime profiles and task-grounded skills, which guide tool selection and program structure. Extensive experiments involving eight models and seven multimodal benchmarks show that unconstrained code generation alone is insufficient and can lead to suboptimal performance. In contrast, equipping the harness with task-grounded skills consistently improves performance over direct model inference in most evaluation settings. As a preliminary exploration of code-driven harnesses, CodeOmni demonstrates the potential of executable code as a shared and extensible substrate for multimodal model interaction, and provides insights toward the development of more capable and general harnesses.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.