ThymeV2: Towards Infinite Visual Exploration with Finite Context
Abstract
Complex perception tasks often require multimodal models to explore and refine visual evidence over multiple interactions. Recent multimodal agents have begun to demonstrate such behavior, yet scaling this capability toward sustained exploration remains challenging. Current systems often operate within limited exploration horizons, and extending exploration naturally causes visual contexts to grow rapidly. In this work we introduce ThymeV2, a multimodal model for sustained active perception in images. ThymeV2 combines adaptive multi-turn visual exploration with an evolving visual memory, allowing the model to acquire new evidence while selectively maintaining useful observations. To unlock this capability, we design visual tools for multi-region image acquisition and removing unnecessary observations, construct large-scale multi-turn exploration dataset and develop a two-stage post-training recipe. We first perform supervised fine-tuning (SFT) on turn-level examples using the visual context available at each turn. Following SFT, our reinforcement learning (RL) recipe combines outcome and process signals with fine-grained control over policy updates for stable optimization. Together, these designs support deep exploration, self-correction, and stable learning over extended interactions. We also introduce StratBenchFinal, a challenging visual exploration benchmark with 893 manually reviewed examples across five difficulty levels. ThymeV2 achieves performance comparable to Gemini 3.1 Pro with the same tools and outperforms Gemini 3.1 Pro without tools by up to 12.76 percentage points on average across six challenging perception benchmarks.We will release our models, data, and code to facilitate future research.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.