When Fine-Grained Credit Assignment Helps and Hurts: Multimodal Agentic Reinforcement Learning through Video Editing
Abstract
Reinforcement learning (RL) has shown promise for training multimodal agents to reason and use tools. However, it remains unclear when fine-grained credit assignment improves learning and when it fails to improve task performance. Video editing provides a useful testbed for this question, combining multimodal perception, long-horizon planning, and tool use with delayed outcome feedback. We construct four movie-trailer editing task types—Content Selection, Narrative Sequencing, Rough Cut, and Fine Cut—that reflect real-world editing workflows and span different levels of task complexity and interaction horizon. Complex, long-horizon tasks benefit more from fine-grained credit assignment, with a 45.4% relative gain over outcome-only GRPO on Fine Cut tasks averaging more than 60 turns under outcome-only GRPO. On simpler tasks averaging 30–40 turns under the same baseline, outcome-only GRPO remains competitive, with fine-grained credit assignment yielding at most a 10.4% relative gain. Process rewards from an LLM judge initially improve tool use and encourage visual inspection. As training progresses, however, GRPO+PRM with a fixed process-reward weight increasingly exhibits excessive reasoning and redundant tool calls: process rewards continue to rise while task completion rates and outcome scores decline. Decaying the process-reward weight to zero mitigates this reward-hacking behavior while preserving the early gains. Despite dense token-level supervision, on-policy self-distillation (OPSD) yields limited and unstable gains and is highly sensitive to the choice of privileged information. Our analysis points to unreliable teacher guidance at student-visited states and dilution of informative signals over long trajectories.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.