OmniPlex-RL: Context Compaction for Native Omni-Modal Agentic Reinforcement Learning
Abstract
Omni-modal agents answer questions from text corpora, knowledge bases, long documents and audio-video streams, reading evidence natively as pixels, pages, frames and sound. Prior work advances what such agents look at, via active perception trained with reinforcement learning (RL), and what they keep, via summaries, folding and textual memory. Yet native looks exhaust the window before the call budget, compaction turns raw evidence into text that cannot be re-opened, and folding leaves no consistent sample or group for credit. To address these challenges, we propose OmniPlex-RL, a native omni-modal agent that compacts its context into segments linked by an omni-modal hypergraph memory whose units keep the call re-opening their raw media. It trains each segment as a sample under single-rollout RL, with a per-modality moving batch baseline. On twelve text, multimodal and omni-modal benchmarks, OmniPlex-RL leads on all twelve, 4.5 F1 above the strongest baseline.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.