Morph-V: Autonomous Growth of Reusable Observation Programs for Long-Video Reasoning
Abstract
Long-video reasoners adapt their observations to individual questions, yet repeatedly plan observation from scratch when similar requirements recur across videos. This prevents successful observation procedures from being retained and reused, limiting the accumulation of observation capability. To this end, we propose MORPH-V, a two-phase framework that couples coverage-guided program acquisition with applicability-guided reuse through a persistent program library. During coverage-guided acquisition, MORPH-V identifies recurring observation requirements unsupported by the current library, composes candidate programs from typed primitives, verifies them structurally without answer outcomes, and stores the verified programs for reuse. Each newly stored program expands library coverage, redirecting subsequent acquisition toward the remaining unsupported requirements. During applicability-guided reuse, MORPH-V selects an applicable program, grounds its observation requirement in the current video, and executes it to construct evidence for the reasoner. Across five video evaluation populations and two frozen reasoners, MORPH-V improves over the corresponding observation controls. Further analysis reveals rising macro-averaged accuracy across reloaded checkpoints and that fixed deployment programs benefit both reasoners.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.