acceptodds
Under review as a conference paper at ICLR 2027

CodeOmni: A Code-Driven Harness for Extensible Omni-Modal Reasoning

Abstract

The code-as-action/tool paradigm has been widely explored in prior work. However, existing approaches either underutilize the expressiveness of code or remain tailored to a single scenario. To address these limitations, we propose CodeOmni, a code-driven harness that uses executable programs to orchestrate interactions between models and harness components, providing a unified execution substrate for image, video, audio, and audio-visual tasks. In CodeOmni, frozen models generate Python programs that jointly support tool orchestration, control flow, exact computation, and task-local state management. A variable-mediated workspace further stores intermediate artifacts outside the main conversational context and exposes them only when needed, reducing unnecessary context accumulation. New tasks can be incorporated through lightweight runtime profiles and task-grounded skills, which guide tool selection and program structure. Extensive experiments involving eight models and seven multimodal benchmarks show that unconstrained code generation alone is insufficient and can lead to suboptimal performance. In contrast, equipping the harness with task-grounded skills consistently improves performance over direct model inference in most evaluation settings. As a preliminary exploration of code-driven harnesses, CodeOmni demonstrates the potential of executable code as a shared and extensible substrate for multimodal model interaction, and provides insights toward the development of more capable and general harnesses.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.