STaR-VL: Capability-Aligned Reinforcement Learning for Visual-Temporal Memory
Abstract
Vision-language models can recognize objects in individual images yet struggle to retain and compare observations across a changing scene. We present STaR-VL, a staged post-training framework for visual-temporal memory on a frozen Qwen2.5-VL backbone. Three LoRA adapters learn per-frame perception, symbolic temporal memory through a recurrent memory module, and transfer to visual histories, respectively; each stage combines supervised fine-tuning with a capability-aligned reinforcement-learning reward. Our contributions are a capability-aligned training framework and a controlled evaluation separating its gains from in-domain joint training, persistent-memory augmentation, and inference-time prompting. On SpaMEM, STaR-VL-3B reaches a V-score of 47.60, compared with 21.85 for the untrained backbone, 40.80 for a joint rank-48 LoRA, and 38.90 for an adapted ReMEmbR system on the same backbone. A structured memory-table prompt raises the untrained model to 22.21. The same trained checkpoints transfer to held-out benchmarks: the 7B model improves VSI-Bench from 33.3 to 46.4, with smaller gains on EgoSchema and TempCompass. Visual-history and symbolic-history controls reveal complementary effects of frame coverage and accessible scene state. These results support staged capability-aligned training for object recall, event detection, and cumulative memory over explicitly ordered observations, while identifying fine-grained grounding and temporal localization as remaining challenges.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.