acceptodds
Under review as a conference paper at ICLR 2027

AVRT: Audio-Visual Reasoning Transfer through Single-Modality Teachers

Abstract

Recent advances in reasoning models have shown remarkable progress in text-based domains, but transferring those capabilities to multimodal settings, e.g., to allow reasoning over audio-visual data, still remains a challenge, i.a., because of the limited availability of high-quality reasoning data in targeted multimodal combinations. To address this problem, we introduce AVRT, a novel framework that generates high-quality audio-visual reasoning traces from single-modality teacher models. We generate independent vision- and audio-reasoning traces via models specialized to reason over their respective modalities and merge the resulting traces with an LLM merger model. The resulting multimodal traces are used in a supervised fine-tuning (SFT) cold start to adapt the target model to audio-visual reasoning traces first, before training it in a second reinforcement learning stage on larger-scale data. Our evaluation shows that the proposed pipeline based on generated multimodal traces for SFT allows models to achieve superior performance on various datasets, i.a., OmniBench and DailyOmni compared to RL alone, establishing a new training pipeline for audio-visual reasoning models and opening ways to train reasoning models on multimodal data.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.